Runbook — the shared devbox
Audience: every coding agent (Claude, Codex, human) working on this machine. This is machine-level knowledge that lives in no single repo's code. Distilled 2026-09-07 from coordinator session memory; verify anything load-bearing against the live box before relying on it.
Layout
/opt/upsquad/<repo>— shared main checkouts. NEVER work in them (CLAUDE.md → Parallel agent dispatches)./opt/upsquad-worktrees/<repo>/<agent>-<issue>/— agent worktrees, created byscripts/agent-worktree.sh. The script is repo-agnostic: run it from inside any repo's main checkout and it operates on that repo (client/admin have no local copy — use core's)./opt/upsquad-keys/— bot-app private keys; the access mechanics are CLAUDE.md → "Bot authentication" (the single copy). One delta the law does not say:sgrunsdash, so POSIX syntax only inside the-cstring (no${var:0:8}, no arrays) — put real logic in a bash script file and invoke that./opt/upsquad/ops/— box-level ops artifacts (e.g.post-reboot-checklist.md).~/.cloudflared/upsquad-access-token.env— Cloudflare Access service token (CF-Access-Client-Id/Secretheaders). The defaultcert.pemin that directory belongs to the WRONG Cloudflare account; always use the upsquad token/cert explicitly.
Stacks running on this box
- Dev stack: 27-service docker compose (as of 2026-09-07), entry via Envoy;
make dev-up. - Prod stack: the SAME box, separate ports/DB/Redis, managed independently (there is no
maketarget for it). Assume both stacks are live at all times — do not stop or restart containers you did not start. - Dev deploys are pull-based: CI publishes images to GHCR and a reconciler pulls them (with rollback). Do not
docker buildimages by hand for dev deploys. Deploy jobs run as GitHub Actions on the self-hosted runner. - Port gotchas: Grafana occupies
:3001. Browser E2E must go through the h2 edge:10443, never:10000. Coder runs atlocalhost:3080(public:coder.upsquad.aibehind CF Access,@upsquad.aionly).
Resource guards (why your job may be refused or killed)
- Disk:
scripts/dev-disk-guard.shruns every 15 min and at ≥85% used trims the Go build cache and then invokes the worktree pruner;agent-worktree.shprints adisk:line on every dispatch and REFUSES at ≥95%. If refused:bash scripts/prune-agent-worktrees.sh --reportanddocker system dfbefore considering the override env. - The Go build cache is the largest wholly-regenerable consumer when it has regrown, and the guard now bounds it (#3508). It reached 19 GB / 70,411 entries with nothing written to it in 48 h, because 182 worktrees share one GOCACHE and Go keys entries partly on build context — the same package compiled in two worktree paths is two entries. It is not always the biggest thing on the box: dated snapshot 2026-09-26T15:52Z,
/opt/upsquad-worktrees21 G, docker images 25.3 G (1.2 G reclaimable) + volumes 10.9 G,/home/vb18 G,/tmp8.6 G, and the cache itself 1.1 MB nine hours after a hand-clear. Re-derive before quoting any of that. Two tiers, both logged to/var/tmp/dev-disk-guard.logwith the bytes freed, and the total also lands in/var/tmp/dev-disk-guard.statusasgocache_freed_mb:- at ≥85%, delete entries untouched for >24 h (
DISK_GUARD_GOCACHE_MAX_AGE_H). Free:markUsed()rewrites an entry's mtime on every cache hit (at most hourly), so nothing in anyone's warm set is older than that. - at ≥88% (
DISK_GUARD_GOCACHE_HARD_PCT) and only if the age tier left the cache over 10 GB (DISK_GUARD_GOCACHE_CAP_MB), evict oldest-first back to the cap, never touching an entry younger than 60 min (DISK_GUARD_GOCACHE_TIER2_FLOOR_MIN). The cap is 5× the 2.1 GB one worktree's full build+test compile measures here. The 60-min floor is the list-then-delete race bound, derived from Go's own 1 hmtimeInterval: tier 2 lists, sorts, thenrms without re-checking mtime, and a-dfile removed afterOutputFile()handed out its path is a build failure. If everything left is inside the floor, the guard logs that it is still over the cap and refuses to widen it. - Kill switch:
DISK_GUARD_GOCACHE_TRIM=0. It never runsgo clean -cache— measured cold-vs-warm on this box is +48 s forgo build ./...and +111 s for a full test compile, per worktree, per firing. - It deletes entry files and never the 256 shard directories. That matters because of the read side, not the write side: a failed cache write is swallowed for plain
go build/go test(buildid.go:722-733returns thePuterror onlyif b.NeedExport), whereas a cache hit is consumed straight out of<cache>/xx/<hash>-d, sogo clean -cacheremoving the shard under a running build is the[build failed]on a~/.cache/go-buildpath. An absent entry file is a plain miss. If you still see a build error naming a path under~/.cache/go-build, that is a finding worth an issue, not expected behaviour. - It only ever trims
$GOCACHE— one cache, the one it resolved. Another user's cache is out of reach: measured 2026-09-26,/home/ashikis 0751 but/home/ashik/.cacheis 0700, so the barrier is the parent and a leaf-level glob cannot even expand. Every run emits anunreclaimed:line for each home it could not see, naming the blocking directory and its mode (that leaf is only 1 MB today — it is an invisibility note, not a size claim). Check those lines before concluding the guard is broken because the disk did not move. - After merging a change to this script you must deploy it by hand. The installed unit runs
/opt/upsquad/ops/dev-disk-guard.sh, not the repo copy, and there is no sync mechanism —deploy_driftonly reports the mismatch into the marker./opt/upsquad/opsisdrwxrwsr-x vb dev, so from a checkout at the merged SHA:sudo install -m 0755 -o vb -g dev scripts/dev-disk-guard.sh /opt/upsquad/ops/dev-disk-guard.sh, then confirm"guard_drift":"none"in/var/tmp/dev-disk-guard.statuson the next 15-min tick. A symlink into a shared checkout is not acceptable — the tree it points at is branch-dependent.
- at ≥85%, delete entries untouched for >24 h (
- Memory:
earlyoomis installed and configured to PREFER killing agent processes (--prefer ^(claude|node)$) while explicitly avoiding infrastructure (postgres, envoy, dockerd, sshd, cloudflared, …). If your session or build vanished under memory pressure, you were most likely the chosen victim — the earlyoom log names its kills; do not go hunting for another memory hog first. Repo-widego test -tags integration ./...has been OOM-killed here more than once — scope integration runs to the packages you touched, one package at a time for the heavy ones.
Secrets
- LLM/provider keys live in the platform vault — never in
.envfiles, shell profiles, or committed config. - The vault is only as private as its passphrase. Before entering a real key on beta-dev,
python3 scripts/vault-rekey.py --checkmust show every consumer on onetargetpassphrase, notdev-literalorempty. Cutover, re-key and rollback:docs/runbooks/vault-passphrase.md(#3476). - Bot GitHub token minting is CLAUDE.md → "Bot authentication" (the single copy). Two deltas from hard experience: call
gh-token.pyby ABSOLUTE path (relative paths break the moment you are not at the repo root), and assert the result is non-empty before using it. - Never
export GH_TOKENin a shell profile.scripts/lib/git-https-fetch.sh(used by worktree creation and pruning) prefers$GH_TOKENfrom the environment over minting a bot token, so a stale profile PAT silently breaks every authenticated fetch withfatal: could not read Username for 'https://github.com'(found 2026-09-07 as the root cause of the #3088 failure class) — and agents must never inherit human tokens anyway. If a helper fails that way, checkenv | grep GH_TOKENfirst and run it withenv -u GH_TOKEN.
Escalation
sudois forbidden for agents. Anything that needs root is a finding to report to the founder, never a thing to route around.