Skip to main content

Runbook — the shared devbox

Audience: every coding agent (Claude, Codex, human) working on this machine. This is machine-level knowledge that lives in no single repo's code. Distilled 2026-09-07 from coordinator session memory; verify anything load-bearing against the live box before relying on it.

Layout​

  • /opt/upsquad/<repo> — shared main checkouts. NEVER work in them (CLAUDE.md → Parallel agent dispatches).
  • /opt/upsquad-worktrees/<repo>/<agent>-<issue>/ — agent worktrees, created by scripts/agent-worktree.sh. The script is repo-agnostic: run it from inside any repo's main checkout and it operates on that repo (client/admin have no local copy — use core's).
  • /opt/upsquad-keys/ — bot-app private keys; the access mechanics are CLAUDE.md → "Bot authentication" (the single copy). One delta the law does not say: sg runs dash, so POSIX syntax only inside the -c string (no ${var:0:8}, no arrays) — put real logic in a bash script file and invoke that.
  • /opt/upsquad/ops/ — box-level ops artifacts (e.g. post-reboot-checklist.md).
  • ~/.cloudflared/upsquad-access-token.env — Cloudflare Access service token (CF-Access-Client-Id/Secret headers). The default cert.pem in that directory belongs to the WRONG Cloudflare account; always use the upsquad token/cert explicitly.

Stacks running on this box​

  • Dev stack: 27-service docker compose (as of 2026-09-07), entry via Envoy; make dev-up.
  • Prod stack: the SAME box, separate ports/DB/Redis, managed independently (there is no make target for it). Assume both stacks are live at all times — do not stop or restart containers you did not start.
  • Dev deploys are pull-based: CI publishes images to GHCR and a reconciler pulls them (with rollback). Do not docker build images by hand for dev deploys. Deploy jobs run as GitHub Actions on the self-hosted runner.
  • Port gotchas: Grafana occupies :3001. Browser E2E must go through the h2 edge :10443, never :10000. Coder runs at localhost:3080 (public: coder.upsquad.ai behind CF Access, @upsquad.ai only).

Resource guards (why your job may be refused or killed)​

  • Disk: scripts/dev-disk-guard.sh runs every 15 min and at ≥85% used trims the Go build cache and then invokes the worktree pruner; agent-worktree.sh prints a disk: line on every dispatch and REFUSES at ≥95%. If refused: bash scripts/prune-agent-worktrees.sh --report and docker system df before considering the override env.
  • The Go build cache is the largest wholly-regenerable consumer when it has regrown, and the guard now bounds it (#3508). It reached 19 GB / 70,411 entries with nothing written to it in 48 h, because 182 worktrees share one GOCACHE and Go keys entries partly on build context — the same package compiled in two worktree paths is two entries. It is not always the biggest thing on the box: dated snapshot 2026-09-26T15:52Z, /opt/upsquad-worktrees 21 G, docker images 25.3 G (1.2 G reclaimable) + volumes 10.9 G, /home/vb 18 G, /tmp 8.6 G, and the cache itself 1.1 MB nine hours after a hand-clear. Re-derive before quoting any of that. Two tiers, both logged to /var/tmp/dev-disk-guard.log with the bytes freed, and the total also lands in /var/tmp/dev-disk-guard.status as gocache_freed_mb:
    • at ≥85%, delete entries untouched for >24 h (DISK_GUARD_GOCACHE_MAX_AGE_H). Free: markUsed() rewrites an entry's mtime on every cache hit (at most hourly), so nothing in anyone's warm set is older than that.
    • at ≥88% (DISK_GUARD_GOCACHE_HARD_PCT) and only if the age tier left the cache over 10 GB (DISK_GUARD_GOCACHE_CAP_MB), evict oldest-first back to the cap, never touching an entry younger than 60 min (DISK_GUARD_GOCACHE_TIER2_FLOOR_MIN). The cap is 5× the 2.1 GB one worktree's full build+test compile measures here. The 60-min floor is the list-then-delete race bound, derived from Go's own 1 h mtimeInterval: tier 2 lists, sorts, then rms without re-checking mtime, and a -d file removed after OutputFile() handed out its path is a build failure. If everything left is inside the floor, the guard logs that it is still over the cap and refuses to widen it.
    • Kill switch: DISK_GUARD_GOCACHE_TRIM=0. It never runs go clean -cache — measured cold-vs-warm on this box is +48 s for go build ./... and +111 s for a full test compile, per worktree, per firing.
    • It deletes entry files and never the 256 shard directories. That matters because of the read side, not the write side: a failed cache write is swallowed for plain go build/go test (buildid.go:722-733 returns the Put error only if b.NeedExport), whereas a cache hit is consumed straight out of <cache>/xx/<hash>-d, so go clean -cache removing the shard under a running build is the [build failed] on a ~/.cache/go-build path. An absent entry file is a plain miss. If you still see a build error naming a path under ~/.cache/go-build, that is a finding worth an issue, not expected behaviour.
    • It only ever trims $GOCACHE — one cache, the one it resolved. Another user's cache is out of reach: measured 2026-09-26, /home/ashik is 0751 but /home/ashik/.cache is 0700, so the barrier is the parent and a leaf-level glob cannot even expand. Every run emits an unreclaimed: line for each home it could not see, naming the blocking directory and its mode (that leaf is only 1 MB today — it is an invisibility note, not a size claim). Check those lines before concluding the guard is broken because the disk did not move.
    • After merging a change to this script you must deploy it by hand. The installed unit runs /opt/upsquad/ops/dev-disk-guard.sh, not the repo copy, and there is no sync mechanism — deploy_drift only reports the mismatch into the marker. /opt/upsquad/ops is drwxrwsr-x vb dev, so from a checkout at the merged SHA: sudo install -m 0755 -o vb -g dev scripts/dev-disk-guard.sh /opt/upsquad/ops/dev-disk-guard.sh, then confirm "guard_drift":"none" in /var/tmp/dev-disk-guard.status on the next 15-min tick. A symlink into a shared checkout is not acceptable — the tree it points at is branch-dependent.
  • Memory: earlyoom is installed and configured to PREFER killing agent processes (--prefer ^(claude|node)$) while explicitly avoiding infrastructure (postgres, envoy, dockerd, sshd, cloudflared, …). If your session or build vanished under memory pressure, you were most likely the chosen victim — the earlyoom log names its kills; do not go hunting for another memory hog first. Repo-wide go test -tags integration ./... has been OOM-killed here more than once — scope integration runs to the packages you touched, one package at a time for the heavy ones.

Secrets​

  • LLM/provider keys live in the platform vault — never in .env files, shell profiles, or committed config.
  • The vault is only as private as its passphrase. Before entering a real key on beta-dev, python3 scripts/vault-rekey.py --check must show every consumer on one target passphrase, not dev-literal or empty. Cutover, re-key and rollback: docs/runbooks/vault-passphrase.md (#3476).
  • Bot GitHub token minting is CLAUDE.md → "Bot authentication" (the single copy). Two deltas from hard experience: call gh-token.py by ABSOLUTE path (relative paths break the moment you are not at the repo root), and assert the result is non-empty before using it.
  • Never export GH_TOKEN in a shell profile. scripts/lib/git-https-fetch.sh (used by worktree creation and pruning) prefers $GH_TOKEN from the environment over minting a bot token, so a stale profile PAT silently breaks every authenticated fetch with fatal: could not read Username for 'https://github.com' (found 2026-09-07 as the root cause of the #3088 failure class) — and agents must never inherit human tokens anyway. If a helper fails that way, check env | grep GH_TOKEN first and run it with env -u GH_TOKEN.

Escalation​

  • sudo is forbidden for agents. Anything that needs root is a finding to report to the founder, never a thing to route around.