Skip to main content

Runbook — upsquad-ci-1, the self-hosted CI runner box

Accepted posture: no HA. One machine, in one datacentre, with no standby. That was a deliberate founder decision (#3242) traded against ~$900/mo of GitHub Actions overage. Everything in this runbook assumes the correct response to "the box is gone" is revert the lanes to hosted runners and carry on, not "restore service under pressure".

  • Issue: #3242
  • Hardware: Hetzner auction #3062626 — Ryzen 7 7700 (8c/16t), 64 GB DDR5, 2 × 1 TB NVMe, HEL1-DC12. €114/mo + €1.90/mo primary IPv4.
  • Hostname: upsquad-ci-1 · IPv4 62.238.87.30 · IPv6 2a01:4f9:3120:18e4::2/64
  • OS: Ubuntu 24.04 LTS, ext4, no RAID (founder-specified). nvme0n1 = OS, nvme1n1 = /data.

1. Access​

Inbound is locked to the devbox (65.21.252.138, 2a01:4f9:c010:e2bd::1) at two independent layers — ufw on the box and the Hetzner Robot edge firewall. There is no public SSH. Root login is disabled entirely.

AccountPurposeKeys
vaisakhfounder, sudo NOPASSWD4 existing devbox keys
ashikengineer, sudo NOPASSWD2 existing keys
ciopsautomation identity, sudo NOPASSWDupsquad-ci-runner
rootlogin disabled (PermitRootLogin no, password locked)—
ssh -i ~/.ssh/id_ed25519_ci_runner ciops@62.238.87.30

All three accounts are password-locked, which is why sudo is NOPASSWD: a key-only account with a locked password can never satisfy a password sudo prompt, so the alternative is not "more secure", it is "sudo does not work".

SSH host key (ed25519): SHA256:gi8Ar9RSDU6kiIsXNz77KP9yJ+I7i0/FQPExHKHzJvU — read off the installed filesystem from the authenticated rescue session before first boot, then confirmed against ssh-keyscan after. If it ever differs, do not accept it.

2. What runs on the box​

UnitWhat it is
upsquad-runner@1..8the runner fleet — one ephemeral GitHub Actions runner per slot
upsquad-runner-1..8.sliceper-slot cgroup: MemoryMax=6G, AllowedCPUs=<n-1>,<n+7>
upsquad-runner.slicefleet-wide cap, MemoryMax=52G
upsquad-registry-cachedocker.io pull-through mirror on :5001, storage /data/registry-cache
upsquad-ci-gc.timerweekly (Sun 04:00) image/cache GC
upsquad-ci-diskguard.timerevery 15 min; >80% used on / or /data escalates to aggressive GC
upsquad-runner-fleet-metrics.timerevery 30s; writes fleet gauges for node_exporter's textfile collector
prometheus-node-exporter:9100, scraped by the devbox Prometheus
earlyoominstalled day 0 — there is no swap on this box
docker-user-firewalldefault-drop for WAN traffic into containers (DOCKER-USER)

Slot anatomy​

Each slot is two containers on a slot-private bridge, both destroyed after every job:

upsquad-dind-N privileged · own dockerd · /var/lib/docker on /data/dind/N (WIPED per job)
upsquad-runner-N UNPRIVILEGED · no host socket · DOCKER_HOST=tcp://dind:2375

The two DinD invariants — do not break these​

Both were established by watching real lanes fail, and both are the sort of thing a later "simplification" removes without understanding.

1. The runner SHARES the sidecar's network namespace (--network container:upsquad-dind-N).

For a non-container job the Actions runner creates services: containers on whatever daemon DOCKER_HOST points at, publishes their ports there, and then expects the job to reach them at localhost. With the sidecar merely routable over a bridge, those ports land in the sidecar's netns and localhost:5432 inside the runner reaches nothing. db.yml — a required check — uses services:.

This grants network reachability only: no filesystem, no PID namespace, no capabilities.

2. The runner and the sidecar share /home/runner/_work AND /tmp, at the same path on both sides.

Any host path a job hands to docker run -v is resolved by the sidecar, in the sidecar's filesystem. If that path does not exist there, Docker helpfully creates an empty directory and mounts it — so the failure surfaces later, as something else.

Found by Migrate Up + Down failing on PR #3244 with:

touch: /etc/pgbouncer/userlist.txt: Read-only file system

scripts/start-test-pgbouncer.sh does CFGDIR="$(mktemp -d)" and bind-mounts it read-only. Note the shape of the failure: not "no such file", but a plausible permissions error three steps downstream. That is what makes this class expensive, and why the fix is the general invariant rather than a TMPDIR special case for one script.

Both _work and /tmp are per-slot directories under /data, wiped between jobs.

Slot N is pinned to physical core N-1 — CPU N-1 and CPU N+7, which are SMT siblings on this Ryzen. The 6 G cap is enforced by the slice, not by per-container --memory, because a slot is two containers and the budget is per slot.

3. Credential isolation — read this before changing a unit​

CI jobs execute branch-controlled code. The design rule is that a job must never be able to reach a credential that outranks it.

  • The minting credential lives at /etc/upsquad-ci/mint.env, 0400 root:root, on the host. It is not bind-mounted into anything. Its two consumers are upsquad-runner-mint-jit.sh and the registration health collector, both run by systemd as root. Every 30s systemd injects it into the collector oneshot's environment. That service uses ProtectSystem=strict, ProtectHome, PrivateTmp, PrivateDevices and an empty CapabilityBoundingSet; it only makes outbound GETs to api.github.com. Its systemctl child receives a minimal environment without the token. The job-to-credential boundary asserted by ci-box-canary.yml is unchanged.
  • No job container gets a host docker socket. Docker for jobs is the per-job dind sidecar.
  • What crosses into the runner container is a JIT config — single-use, and bound to one runner name. A registration token would have been strictly worse: it can enrol arbitrary additional runners.
  • The JIT config arrives on a slot-private bind mount and the entrypoint unlinks it before exec'ing the listener, so it is not in /proc/1/environ and not on disk.

ci-box-canary.yml asserts all of this daily. If you change upsquad-runner-slot.sh, expect that lane to tell you.

Residual risk, stated plainly​

The dind sidecar is privileged. A job that pivots into it can reach the host, and therefore mint.env. That is why the minting credential must be minimal scope — it is the last line of defence, not the first. A self-hosted-runner administration only fine-grained PAT limited to upsquad-ai/upsquad-core bounds an escape to "can enrol runners on one repo". An App key with contents:write would bound it to "can rewrite the product". Never put a /opt/upsquad-keys/ App key on this box.

The unprivileged/privileged split is still worth having: it means a job has to break two boundaries rather than zero.

4. Common operations​

# Fleet state at a glance
systemctl list-units 'upsquad-runner@*'
curl -s localhost:9100/metrics | grep upsquad_runner

# Why is a slot not coming up?
sudo journalctl -u upsquad-runner@3 -n 100 --no-pager

# Reclaim disk right now
sudo /usr/local/sbin/upsquad-ci-gc.sh --aggressive

Restarting: use the drain wrapper, never systemctl restart​

sudo /usr/local/sbin/upsquad-runner-restart.sh # whole fleet, drained
sudo /usr/local/sbin/upsquad-runner-restart.sh 3 7 # just those slots

sudo systemctl restart 'upsquad-runner@*' KILLS RUNNING JOBS. It stops the container immediately and GitHub reports ##[error]The runner has received a shutdown signal.

Not hypothetical: rebuilding the image for an unrelated fix during #3242 took out four in-flight lanes on PR #3250 at once — Go lint, Golden Flow E2E, North-Star SDLC E2E and Proto Generation & Go Compilation.

They did not look like a restart. The proto one surfaced as exit code 130 next to an "out of date — run make proto" message from a step that never finished: a completely convincing, completely wrong diagnosis, sitting there waiting to be acted on. Check for the shutdown-signal line before believing any fleet failure.

The wrapper waits for each slot to go idle first. Slots are ephemeral — one job, then exit — so waiting is all that is ever needed; a restart never has to interrupt anything.

Its oracle for "busy" is a Runner.Worker process inside the runner container, which is a direct observation of work in progress. systemctl is-active is not the oracle: a slot reads active from the moment its container starts, job or no job. (Same trap in a different shape as the dev-box runner that reported online to the API while accepting nothing — core#3212.)

TIMEOUT=0 overrides the wait and interrupts the job. It exists for a genuinely wedged slot; it is not the normal path.

Re-register the fleet​

Registrations are re-minted on every job; there is no long-lived registration to repair. If GitHub's roster is out of sync (stale offline entries under hetzner-ci-1-slot-*), the mint script reaps a same-named runner before it mints, so the fix is simply to restart the slot.

If runners never appear at all, the credential is the first thing to check:

sudo journalctl -u 'upsquad-runner@*' --since -10min | grep -i mint

no minting credential ... and no spooled JIT config means mint.env is absent or its token has expired/been revoked.

5. Rollback to hosted runners​

This is the primary incident response, not a last resort.

Every lane migrated to the box moved by a single-line runs-on change. Reverting that PR returns the lane to ubuntu-latest with no other coupling — no secrets, no cache keys, no image references differ between the two.

5.1 The revert PR heals itself — it does NOT need the box​

The obvious fear is circular: after both waves, 10 of the 11 required contexts run on this box, so a revert PR appears to need the very box it is trying to stop using. With the box dead those jobs queue rather than fail, the queue is ALLGREEN with a 3600 s checkResponseTimeout, and entries are dropped after an hour rather than merged.

That reasoning is wrong, and the reason matters. A pull request's checks run the workflow definitions in its own head ref, not main's. A revert PR sets runs-on: ubuntu-latest in its head ref, so its own required lanes run on GitHub-hosted runners and report normally.

Measured — and the measurement is a clean crossed experiment, not an inference. #3244 migrated go-test but not go-lint; #3250 migrated go-lint but not go-test. Both target main, and main said ubuntu-latest for both at the time:

PR (base main)Go Unit Tests ran onGo lint ran on
#3244 — migrates go-test onlyhetzner-ci-1-slot-4GitHub-hosted
#3250 — migrates go-lint onlyGitHub-hostedhetzner-ci-1-slot-5

The same two lanes land on opposite runner types purely according to which PR's head ref edited which file. Reverting runs-on therefore moves that PR's own checks back to hosted runners.

So the rollback is an ordinary PR: open it, get a review, enqueue, done. No admin powers, no protection changes, no queue bypass.

git checkout -b revert/hetzner-lanes origin/main
git revert --no-edit <wave-b-merge-sha> <wave-a-merge-sha>
git push origin revert/hetzner-lanes
gh api repos/upsquad-ai/upsquad-core/pulls --method POST \
-f title='revert: return CI lanes to hosted runners' \
-f head='revert/hetzner-lanes' -f base='main' \
-f body='upsquad-ci-1 unavailable. Returning lanes to ubuntu-latest.'

The merge queue behaves the same way, and that is observed, not assumed. PR #3211 (e093188dc) is the commit that added the merge_group: trigger to migration-ceiling.yml. Its own queue entry ran Migration Ceiling as a merge_group job — at a moment when main had no such trigger, so the only tree in which that trigger existed was the entry's merged tree. GitHub fired it anyway. The queue therefore evaluates workflow definitions from the entry's merged tree, exactly as pull_request evaluates them from the head ref. Same chicken-and-egg shape as the crossed experiment above, and direct observation rather than documentation.

One limit, stated because it is easy to miss:

  • Other PRs open during the outage stay stuck. They still carry box-bound lanes in their own head refs. A push to main does not re-trigger a PR's checks on its own, so each one needs an explicit rebase:

    gh api repos/upsquad-ai/upsquad-core/pulls/N/update-branch -X PUT

    The revert landing does not unblock them by itself.

    If that returns 422 expected head sha didn't match current head ref, the PR's head has moved (someone pushed — possibly you, moments earlier) and GitHub is refusing an optimistic-concurrency check it was not given. Re-read the head and pass it. Hit for real while landing this very section:

    SHA=$(gh api repos/upsquad-ai/upsquad-core/pulls/N --jq '.head.sha')
    gh api repos/upsquad-ai/upsquad-core/pulls/N/update-branch -X PUT \
    -f expected_head_sha="$SHA"

temporal-tests was never migrated — it is timing-sensitive and stays hosted permanently.

5.2 Break-glass — for a change that is NOT a runs-on revert​

Needed only when something else must land during an outage, so §5.1's self-healing property does not apply. CLAUDE.md sanctions this shape for legitimate admin operations. devops-engineer holds the only administration bit.

Paths are written out in full, with no $B-style variable. An earlier draft of this block used one, set it to the …/protection root, and then appended /protection/enforce_admins to it anyway — every call 404'd with {"message":"Branch not found"}. For a five-line procedure someone will paste under stress, a variable buys nothing and costs exactly that. Verified live, 2026-09-08:

…/branches/main/protection/protection/enforce_admins -> 404 Branch not found
…/branches/main/protection/enforce_admins -> 200 {"enabled":true}

Precondition — run this with the merge queue EMPTY. A direct merge moves main, which forces every in-flight queue entry to rebuild: you take other people's CI with you. On 2026-09-08 the queue did not empty once across a 22-minute poll, so this needs announcing and scheduling, not opportunism.

gh api graphql -f query='{repository(owner:"upsquad-ai",name:"upsquad-core")
{mergeQueue(branch:"main"){entries(first:20){nodes{position pullRequest{number}}}}}}'
# 1. SNAPSHOT FIRST, and arm the restore before anything is opened.
gh api repos/upsquad-ai/upsquad-core/branches/main/protection > /tmp/prot-before.json
trap 'gh api repos/upsquad-ai/upsquad-core/branches/main/protection/enforce_admins -X POST' EXIT

# 2. Open the glass.
gh api repos/upsquad-ai/upsquad-core/branches/main/protection/enforce_admins -X DELETE

# 3. Merge. Substitute the real PR number for N.
gh api repos/upsquad-ai/upsquad-core/pulls/N/merge -X PUT -f merge_method=squash

# 4. Restore explicitly (the trap is the backstop, not the plan), then PROVE it.
gh api repos/upsquad-ai/upsquad-core/branches/main/protection/enforce_admins -X POST
gh api repos/upsquad-ai/upsquad-core/branches/main/protection > /tmp/prot-after.json
diff <(jq -S . /tmp/prot-before.json) <(jq -S . /tmp/prot-after.json) # must be empty

The trap is deliberate: step 4 is the step that must not be skipped, and a shell that dies between 2 and 4 leaves main unprotected. A window left open is a worse outage than the one being fixed.

Rehearsed 2026-09-08 on a scratch branch protected exactly like main (enforce_admins: true, 1 required review left deliberately unsatisfied, and a required context Phantom Check (never reports) that no workflow reports — the box-is-dead condition):

stepresult
control merge, glass closed405 — "At least 1 approving review is required… Required status check "Phantom Check (never reports)" is expected."
DELETE …/enforce_adminsenabled=false; review + context requirements still configured — only admin enforcement lifts
merge, glass open{"merged": true} — past both an unsatisfied review and a never-reporting check
POST …/enforce_adminsenabled=true; full protection read-back identical
glass-open window15 s (API calls ~3.5 s of it)

The control is the load-bearing half. Without it a successful merge would prove nothing — the PR might have been mergeable anyway.

The drill used the same endpoints written above, against the drill branch instead of main (…/branches/<branch>/protection/enforce_admins). If you change the snippet, re-run the drill; the whole point of this section is that what is written down is what was executed.

The merge-queue leg is now rehearsed too — on main itself, 2026-09-08, founder-supervised (core#3287). It could not be done on a scratch branch: a scratch branch cannot have a merge queue, and the config is UI-only and per-branch. There is no safe probe either — GitHub evaluates mergeability and draft status before protection, so a conflicting PR returns only "Pull Request has merge conflicts" (#3277) and a draft only "Pull Request is still a draft" (#3280). The only test was a real merge on main.

enforce_admins: false DOES lift the merge-queue gate.

stepresult
control merge, glass closed405 naming three gates — see below
DELETE …/enforce_adminsenabled=false
merge, glass open{"merged": true} — past all three
POST …/enforce_adminsenabled=true; protection read-back identical to the snapshot
merge-queue configuration afteridentical (checked separately — see the warning below)
main tree after scratch commit + revertidentical tree SHA to pre-drill
glass-open window3.9 s merge, 3.8 s revert

It is a bigger hammer than "lifts the queue". Know this before you use it.

The control refusal named three gates, not one:

405 Changes must be made through the merge queue
At least 1 approving review is required by reviewers with write access.
11 of 11 required status checks have not succeeded: 6 expected.

With the glass open, that same PR merged at 11:41:51Z carrying zero approving reviews (reviews: []) and 4 of its 11 required contexts unsucceeded — and three of those four had not even started: Compose Smoke Test (started 11:45:22Z), Go Vet (integration tag) (11:44:40Z) and Migrate Up + Down (11:47:18Z) all began after the merge; only Bash script tests was running at the time. enforce_admins: false yields the queue, the review requirement, and the status-check requirement together.

(Derive that from check-run completed_at and started_at against merged_at — not from the control's 405, which said "6 expected" five minutes earlier and describes a different moment. Note the method's limit: completed_at supports "unsucceeded", not "running"; started_at is what separates those. Check-runs expose no created_at and actions/jobs/{id} 403s for bot tokens, so "queued" versus "never created" is not answerable at our permission level — do not claim it.)

That is the right capability for a box-is-dead recovery and the reason this is never a routine tool. Nothing about the queue makes it safer than the other two gates it drops on the way past.

If the merge still returns "Changes must be made through the merge queue" with the glass already open, the behaviour has changed since it was measured. Restore enforce_admins first, then enqueue normally — an ordinary PR is almost always the answer, because §5.1 means a runs-on revert does not need this section at all. Re-open the question with a fresh supervised drill rather than improvising; core#3287 is the template.

Do NOT reach for "uncheck Require merge queue in the UI". That was the old fallback in this section and it is no longer needed — but it is still a trap if someone tries it: unchecking discards the queue configuration, and there is no REST endpoint and no GraphQL mutation to put it back. If you ever do need it, record the config first with the GraphQL query in landing-a-pr.md §6. The drill snapshotted it before opening anything, for exactly this reason, and confirmed it unchanged afterwards.

6. Box-dead playbook​

  1. Confirm it is the box, not GitHub. Are hosted lanes green while only hetzner-ci lanes queue? CIBoxDown firing while the devbox scrapes everything else fine is the same signal.
  2. Do not wait. Queued self-hosted jobs do not fail fast — they sit. Open the revert PR (§5.1) first, then investigate. It is an ordinary PR and needs no admin powers: its own lanes run hosted because it is the thing that sets them back to ubuntu-latest. If something other than a runs-on revert has to land during the outage, that is §5.2 — now rehearsed on main end to end (core#3287), including the merge-queue gate. Note what it costs: it drops the review requirement and every status check along with the queue.
  3. Hetzner Robot → reset / rescue. The box is reprovisionable from scratch in about 30 minutes; nothing on it is precious. /data holds only caches, all of which are re-fetchable.
  4. On the way back up, the two layers that will lock you out if you forget them: the Robot edge firewall (allowlists the devbox only) and ufw. From anywhere other than the devbox, both look like "the box is down".

7. Deliberate design choices that look like bugs​

  • node_exporter is a host package, not a container. Anything Docker publishes is DNAT'd and filtered in FORWARD, which ufw's inbound policy never sees — a containerised exporter would not actually be governed by the "9100 from the devbox only" rule.
  • The dind store is wiped every job. That costs a re-pull, which is why the host runs a docker.io pull-through mirror: the re-pull comes off local NVMe. Persisting it would be a cross-job channel for planting image tags.
  • GOMODCACHE is shared across slots; GOCACHE is not. Module-cache entries are checksum-verified against go.sum/the sumdb, so poisoning is detected. Build-cache entries are keyed by an action-ID hash and are not independently verified, so a shared GOCACHE is a genuine cross-job code-execution channel. Per-slot keeps the warm-cache win and cuts the blast radius from eight slots to the next job on one.
  • No swap. Deliberate; earlyoom is the mechanism instead. A CI box that swaps does not fail, it hangs.
  • Unattended-Upgrade::Automatic-Reboot "false". Jobs run at arbitrary hours and there is no failover. Reboots are a manual step here.

8. Cache-corruption incident — diagnostics first (#3312)​

Symptom. Image-building jobs — Compose Smoke Test, Container Image Scan, Security Scan, Wave 2 Smoke, and the DB Migrations image steps — fail with:

short read: expected 32 bytes but got 0: unexpected EOF

Recurs across different slots and regenerates within hours of a full docker builder prune -af, so the cause is generative, not a stale artefact. The box has ample free disk — not disk exhaustion.

What the error actually is (read this before "fixing" a store)​

It is not buildkit and not "a zero-length file where a digest blob should be". It is containerd content.Copy() (core/content/helpers.go:221) during a blob ingest:

if size != 0 && copied < size-ws.Offset {
return fmt.Errorf("short read: expected %d bytes but got %d: %w",
size-ws.Offset, copied, io.ErrUnexpectedEOF)

expected = size - ws.Offset = the bytes remaining on a (likely resumed) transfer; got = what the source reader delivered mid-stream. So the source — for docker.io images, the upsquad-registry-cache pull-through mirror (:5001, /data/registry-cache, §2/§7) — served a short body during a fetch. The string says nothing about a committed corrupt blob on disk.

Why the on-disk store is UNCONFIRMED​

distribution publishes a blob to blobs/sha256/xx/<digest>/data only via validateBlob(digest) → moveBlob (atomic rename). A truncated, raced, or OOM-killed ingest fails validation and is cancelled, leaving an orphan under _uploads/, not a committed short blob. distribution's proxy store also serialises same-digest fetches (inflight map), so a concurrent-pull race does not obviously commit a bad blob either. Therefore we do not know which store (if any) persists a corrupt artefact: it may be transient-in-ingest, in a slot's containerd content store, or in the dind build cache.

The slot-private / wiped-per-job premise (§2 L64, §7) is UNVERIFIED on the box and is now known to be partly stale. A box measurement 2026-09-14 (#3312 issuecomment-5659325634) found §2 already wrong post-expansion — 12 slots, not 8; MemoryMax 4G, not 6G; the fleet-wide cap gone — and that a read-only ciops SSH path exists for diagnostics. So do not treat "the build cache is wiped, therefore the mirror is the store" as settled: the prune -af bought ~4h datum cuts against a pure-mirror story (a host prune reaches nothing the per-job dinds use), which means either the 4h was coincidence or the build cache is back in frame. No candidate is ruled out on the runbook alone. Run §8.2 to get ground truth.

Two root-cause candidates — both open until §8.2​

  1. OOM-kill mid-write. No swap; earlyoom is the mechanism (§7). A writer killed mid-ingest leaves a cancelled _uploads/ orphan and a short served body on a resumed fetch — matching the error.
  2. Concurrent-pull race on the shared mirror. Multiple slots fetching the same layer through :5001 racing the mirror's serve/ingest.

8.1 The diagnostic (read-only by default)​

deploy/ci-runner/upsquad-ci-cache-guard.sh is a read-only diagnostic — it reports, it does not delete, unless you pass --prune.

  • Primary — content-hash verify. Walks blobs/sha256/.../data and reports every blob whose sha256(content) != digest dirname. That is corruption by definition in a content-addressed store, whatever the upstream cause. Scoped by default to blobs modified in the last --hours (default 6); --full hashes the whole mirror.
  • Secondary — zero-length sweep. Scoped to blobs/sha256, -mmin +10 (never an in-flight commit), excluding the legitimate empty digest e3b0c442…. Expected to find nothing — its presence would itself be a finding.
  • _uploads/ orphan summary — count + age of incomplete ingests, the expected residue of a killed/raced transfer.

--prune (opt-in) removes what the report finds and runs docker builder prune -f --filter until=1h (a no-op by design on the host — build caches live in the dinds — kept only in case §8.2 shows the premise stale). --dry-run pairs with --prune to preview.

The systemd .service/.timer ship in deploy/ci-runner/ but the service's ExecStart is the read-only scan (no --prune), and the timer must stay DISABLED until §8.2 names the store. Installing a deleter on an unconfirmed target risks the next quiet day being read as containment.

Install the diagnostic (founder / vaisakh / ciops read-only path, on the box)​

sudo install -m 0755 deploy/ci-runner/upsquad-ci-cache-guard.sh \
/usr/local/sbin/upsquad-ci-cache-guard.sh
sudo install -m 0644 deploy/ci-runner/upsquad-ci-cache-guard.service \
/etc/systemd/system/upsquad-ci-cache-guard.service
sudo install -m 0644 deploy/ci-runner/upsquad-ci-cache-guard.timer \
/etc/systemd/system/upsquad-ci-cache-guard.timer
sudo systemctl daemon-reload
# Do NOT `enable --now` the timer yet — run §8.2 first.

Rollback​

sudo systemctl disable --now upsquad-ci-cache-guard.timer 2>/dev/null || true
sudo rm -f /etc/systemd/system/upsquad-ci-cache-guard.service \
/etc/systemd/system/upsquad-ci-cache-guard.timer \
/usr/local/sbin/upsquad-ci-cache-guard.sh
sudo systemctl daemon-reload

8.2 Get ground truth in one box login​

Run during or right after a fresh ejection; correlate timestamps with the failing queue run. All of these are read-only.

# (1) Is the mirror's on-disk store ACTUALLY corrupt? The load-bearing check —
# settles whether a persisted corrupt blob exists at all.
sudo /usr/local/sbin/upsquad-ci-cache-guard.sh --full # report-only
# Any "CORRUPT (digest mismatch)" or "FINDING (zero-length…)" line => the
# mirror persists the artefact => enable the timer / run one --prune.
# Clean, plus _uploads orphans present => corruption is TRANSIENT-in-ingest
# (candidate 1 or 2), NOT a persisted blob => the deleter would be inert;
# do NOT enable it. Fix the ingest, not the store.

# (2) _uploads orphans — the expected killed/raced residue (count + age).
sudo find /data/registry-cache/docker/registry/v2/repositories \
-type f -path '*/_uploads/*/data' -printf '%TY-%Tm-%Td %TH:%TM %s %p\n'

# (3) The mirror's own view of the failing digest + short serve.
docker logs --since 6h upsquad-registry-cache 2>&1 | grep -iE 'sha256|short|error|EOF'
# Confirm the mirror's proxy TTL and the dind mirror wiring:
docker exec upsquad-registry-cache cat /etc/docker/registry/config.yml | grep -iA3 proxy
sudo cat /data/dind/1/daemon.json 2>/dev/null | grep -i registry-mirrors # any slot N

# (4) OOM evidence — candidate 1.
journalctl -k --since -6h | grep -iE 'killed process|out of memory|oom-kill'
journalctl -u earlyoom --since -6h --no-pager | grep -iE 'killing|SIGKILL'
# earlyoom naming dockerd/containerd/registry/docker near the ejection =>
# candidate 1. Keep the no-swap posture (§7 — a swapping CI box hangs);
# durable fix is OOMScoreAdjust protecting the store writers. Architect call,
# box-verified.

# (5) Concurrent-pull race — candidate 2.
docker events --filter event=pull --since 10m # during a burst
# 2+ slots on the same digest interleaved with a registry error on it =>
# candidate 2. Durable fix: registry write-locking / per-slot pull namespace.
# Architect call, box-verified.

# (6) Is the dind store still wiped? (§2 is stale — verify, don't assume.)
sudo du -sh /data/dind/* ; sudo find /data/dind -maxdepth 3 -name '*.ingest*' 2>/dev/null
# Compare across two consecutive jobs on the same slot.

Only after (1) shows a persisted corrupt blob do you enable the timer or run sudo upsquad-ci-cache-guard.sh --prune. If (1) is clean, the artefact is elsewhere/transient — take the fix to principal-architect with the (2)–(6) evidence rather than installing a guard that cannot fire.

Registration health collector (#3324)​

The eight-slot fleet is registered to upsquad-ai/upsquad-core, not the organization. Query /repos/upsquad-ai/upsquad-core/actions/runners and match exact names hetzner-ci-1-slot-1 through hetzner-ci-1-slot-8; exclude the separate devbox runner. Do not interpret an empty organization inventory as a fleet outage or silently change registration scope.

check-ci-runner-registration.py replaces the former container-count writer. The existing 30-second timer invokes the checked-in service, which reads the existing root-only runner.env and mint.env. No credential is copied into a runner container. The token needs runner-inventory read access in the same repository it already administers for JIT registration. At one page per poll, this adds approximately 120 API requests/hour.

  • upsquad_runner_slot_registered: the expected name exists in GitHub.
  • upsquad_runner_slot_online and the compatible slot_up: exactly one matching registration is online; busy online runners count as healthy.
  • upsquad_runner_fleet_online: sum of those eight online slot gauges.
  • upsquad_runner_unit_active: independent systemd state, never the fleet count.
  • upsquad_runner_registration_credential_rejected: an explicit authentication or non-rate-limit permission rejection of the shared mint credential.
  • upsquad_runner_registration_api_success / unit_query_success: whether each observation succeeded. Incomplete pagination, malformed responses, redirects and API failures are unknown, not a zero fleet or cached green.

CIBoxRunnerRegistrationMissing warns when an active unit has no unique online registration for five minutes. This accommodates normal ephemeral teardown and JIT registration. CIBoxRunnerRegistrationMetricsUnavailable warns after two minutes of query failure or a missing collector. A separate critical CIBoxRunnerCredentialRejected alert covers a sustained 401 or non-rate-limit 403: the same credential mints every ephemeral slot, so rejection threatens re-registration as current jobs finish. Rate limits/timeouts remain unknown. Existing file-mtime monitoring detects a stopped writer. CIBoxFleetDegraded now warns for any sustained loss (1–7 online out of 8 for five minutes); zero has its own alert.

Installation from the reviewed checkout, on the CI host (preserve the current script/service and metric file first for rollback):

# The writer has no CAP_DAC_OVERRIDE: this directory must be root-writable.
stat -c '%U:%G %a' /var/lib/node_exporter/textfile
set -e # Stop on any failed install or service start.
sudo install -D -m 0755 scripts/check-ci-runner-registration.py /usr/local/lib/upsquad-ci/check-ci-runner-registration.py
sudo install -m 0644 scripts/systemd/upsquad-runner-fleet-metrics.service /etc/systemd/system/upsquad-runner-fleet-metrics.service
sudo systemctl daemon-reload
sudo systemctl start upsquad-runner-fleet-metrics.service
sudo systemctl is-active upsquad-runner-fleet-metrics.timer
curl -fsS localhost:9100/metrics | grep '^upsquad_runner_registration_api_success 1$'

Deploy dev/prometheus/rules/ci-box.yml to the Prometheus rules directory and reload Prometheus; verify the rules through /api/v1/rules and compare the scraped fleet gauge with a fresh repository runner inventory. Restoring the saved service and running it rolls the collector back, but also restores the old container-count blind spot; restore the previous rules with it. No runner restart is needed for monitoring deployment.

For recovery, inspect the affected journals and API busy state first. A transport cancellation immediately after successful completion is not proof of a token failure. If a restart is necessary, prevent automatic recycling with a temporary runtime Restart=no drop-in, wait for the ephemeral job to finish and the unit to deactivate, remove that drop-in, reload systemd and restart only that unit. A same-name entry may be removed only once proved dead; the existing JIT mint helper already performs same-name cleanup before minting. Never delete a busy registration or stop its unit to repair a dashboard count.

Verification:

python3 -m unittest tests.scripts.test_ci_runner_registration -v
bash tests/scripts/test_ci_runner_registration_alerts.sh

Deferred scope is tracked in #3468: a safely idle-only recovery watchdog and an authoritative derived fleet roster. This collector alerts; it never restarts a slot. Until that follow-up lands, 8 is encoded in range(1, 9) in the collector, the < 8 fleet rule, and its /8 / of 8 annotations (plus their regression expectations); a resize must change all of these together. Staggered normal recycles can also keep the aggregate below eight while each individual slot recovers promptly; measure that before changing aggregate grace semantics. The per-slot alert evaluates each slot's continuous loss independently.

Grace-period evidence (2026-09-19, 08:00Z onward): 91 complete observed worker-completion → next-listening intervals across slots 1–8 ranged from 28.2s to 210.4s, including the intentional drain of slots 4/5/8. Five minutes exceeds every interval in that sample. This does not bound interrupted/cancelled jobs or GitHub inventory propagation, and does not prove aggregate staggered recycles cannot warn; retain per-slot evidence when investigating a warning.