CI registry corruption recovery
Owner: devops-engineer. Tracking: core#3432; affected PR core#3431, run 35015944268. This is an operator procedure, not a record of completed repair.
Evidence and scope
On 2026-09-16 at approximately 09:26 UTC the mirror returned HTTP 200 and Content-Length 32 for library/golang blob sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1, but the body was empty (SHA-256 e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855). The new probe reproduced this. Direct Docker Hub retrieval returned the correct 32 bytes and expected digest. This is ongoing mirror corruption, not merely a historical cache problem.
Affected run attempt 2 failed Compose Smoke Test on slot 4. The fleet still had 12 active units at inspection. Discover the fleet at execution time; do not assume eight slots from the older runbook. Resizing remains separate (#3312).
Root authority is required. AGENTS.md prohibits agents using sudo; an explicit founder exception or a human operator is required before executing host changes. Do not use a privileged CI job to bypass that boundary.
1. Inventory and drain (operator with root authority)
Record UTC time, unit enablement, Restart settings, CPU/memory properties, images, runner registrations/busy state, DinD names/mounts, builder drivers and volumes. Do not print runner environment files or minting credentials.
Read /usr/local/sbin/upsquad-runner-slot.sh first: confirm ephemeral one-job exit and what cleanup removes. Current documentation says /data/dind/N is wiped after each job; privileged inspection must confirm it. Check for shared buildx volumes.
For every discovered upsquad-runner@N.service, create a uniquely named runtime drop-in /run/systemd/system/upsquad-runner@N.service.d/90-registry-recovery.conf:
[Service]
Restart=no
Run systemctl daemon-reload, then verify Restart=no on EVERY inventoried unit. This prevents respawning after the current ephemeral runner exits; it does not stop any container or cancel a job. A listener that has not yet accepted its one job may accept one final job. Wait for every unit to exit naturally, including ExecStopPost cleanup. Do not stop/restart active runner units or Docker itself.
If an idle listener never receives a job, leave it alone and arrange a normal harmless job through the existing workflow mechanism. Do not replace the wait with a timeout that kills the listener. Record the terminal jobs and confirm no Runner.Worker/build processes remain. Units must be inactive/failed with no start/stop jobs pending before storage maintenance.
If cleanup destroys each sidecar and wipes its private Docker store, record that as a full cache reset for that slot (pruning a vanished daemon is meaningless). For any retained, verified idle sidecar, run and record:
docker exec upsquad-dind-N docker builder prune -af
docker exec upsquad-dind-N docker buildx ls
If a non-default docker-container builder is used, prune that builder explicitly with docker buildx prune --builder NAME -af in its owning daemon, or remove only its positively identified cache volume once all users are stopped. Do not run global docker system/volume prune. Keep a per-slot coverage table: slot, daemon, cache backing store, prune result or verified cleanup removal. No uncovered slot.
2. Reset mirror storage
With the fleet fully drained, verify the live mirror container mounts exactly /data/registry-cache at /var/lib/registry. Stop only upsquad-registry-cache.service. Verify its container is gone and no process/container still mounts that directory. Quarantine the storage on the same filesystem instead of deleting evidence:
set -eu
stamp=$(date -u +%Y%m%dT%H%M%SZ)
quarantine="/data/registry-cache.corrupt-${stamp}"
test ! -e "$quarantine"
test -d /data/registry-cache
test ! -L /data/registry-cache
mv -- /data/registry-cache "$quarantine"
install -d -m 0755 /data/registry-cache
chown --reference="$quarantine" /data/registry-cache
chmod --reference="$quarantine" /data/registry-cache
systemctl start upsquad-registry-cache.service
python3 /path/to/reviewed/scripts/check-ci-registry-blob.py
The new directory preserves the old ownership/mode. Starting the mirror refills from upstream. Require exit 0 from the digest probe. On failure leave runners drained and inspect mirror logs and upstream reachability. Never silently restore known-corrupt storage. After successful recovery, deletion of the recorded quarantine is a separate explicit cleanup decision; it costs no extra copy space while diagnosing.
3. Clean build and resume
Use one drained slot's verified preparation path to create its DinD sidecar without starting its GitHub listener. Inspect the live helper before invoking it; do not invent helper options. Alternatively use a disposable verification sidecar with the same pinned image, network/mirror configuration and resource limits. Build inside that slot's normal client environment with the source at 8a99a3eb879e557b02feab6a9d29ba614b11d601 (the failed run's head):
docker compose -f docker-compose.dev.yml build --no-cache agent-orchestrator
Record slot, source SHA, daemon ID, image digest, build duration and exit status. Do not put bot credentials or the host Docker socket in the build container. The normal workflow sets COMPOSE_FILE=docker-compose.dev.yml and builds agent-orchestrator plus agent-worker; use its environment if additional variables are needed. Remove the temporary verification resources after recording evidence.
Only once the mirror probe and clean build pass, remove ONLY the recovery drop-ins created above, daemon-reload, verify original Restart settings, and start the originally active slots. Preserve their images, CPU pins, memory limits and enablement. Verify online registrations and completed real jobs.
With the devops bot, request the user-authorized rerun:
gh api -X POST repos/upsquad-ai/upsquad-core/actions/runs/35015944268/rerun-failed-jobs
gh api 'repos/upsquad-ai/upsquad-core/actions/runs/35015944268/jobs?per_page=100' \
--paginate --jq '.jobs[] | [.name,.status,.conclusion,.runner_name,.html_url] | @tsv'
Require a new attempt with Compose Smoke Test completed/success and verify its build step actually ran. Do not claim success from a skipped job or old attempt.
4. Install the probe
From the reviewed checkout on the CI host, as the authorized operator:
install -d -m 0755 /usr/local/lib/upsquad-ci
install -m 0644 scripts/check-ci-registry-blob.py /usr/local/lib/upsquad-ci/
install -m 0644 scripts/systemd/upsquad-registry-probe.service /etc/systemd/system/
install -m 0644 scripts/systemd/upsquad-registry-probe.timer /etc/systemd/system/
systemctl daemon-reload
systemctl start upsquad-registry-probe.service
# Refuse to overwrite any pre-existing collector entry; inspect it instead.
test ! -e /var/lib/node_exporter/textfile/upsquad_registry.prom
test ! -L /var/lib/node_exporter/textfile/upsquad_registry.prom
ln -s /var/lib/upsquad-registry-probe/registry.prom \
/var/lib/node_exporter/textfile/upsquad_registry.prom
systemctl enable --now upsquad-registry-probe.timer
curl -fsS http://127.0.0.1:9100/metrics | grep upsquad_registry_blob
The service runs as the existing prometheus user, has no Docker/root credential, and writes atomically to its own StateDirectory. The collector symlink follows atomic replacements. The systemd 20-second deadline bounds slow responses.
Deploy dev/prometheus/rules/ci-registry.yml through the existing dev Prometheus configuration/reload procedure. Its configured rules/.yml glob loads the file. Verify the live /api/v1/rules contains both CIBoxRegistryBlobInvalid and CIBoxRegistryProbeStale and /api/v1/query sees fresh upsquad_registry_blob_ series with job="ci-box". A checked-in file alone is not deployed monitoring. If the ci-box scrape is missing, record that blocker; absence of alerts is not proof.
Check twice at least 60 seconds apart: digest result 1 and advancing timestamp. The invalid rule fires after 2 minutes of failure; the stale rule catches a probe older than 180 seconds or missing on a reachable ci-box target after 2 minutes. This probes a known canary, not every blob; other corruption can still require storage reset. Redirects fail deliberately: this local filesystem mirror must serve the bytes itself, not send the probe straight to a healthy upstream.
Verification and audit
python3 -m unittest discover -s tests/scripts -p test_check_ci_registry_blob.py -v
bash tests/scripts/test_ci_registry_alerts.sh
systemd-analyze verify scripts/systemd/upsquad-registry-probe.service scripts/systemd/upsquad-registry-probe.timer
Record on #3431: UTC start/end, operator, per-slot cache coverage, mirror storage quarantine path, before/after blob hashes, clean build evidence, rerun URL/attempt/ job conclusion, installed probe commit and live alert/metric checks. Keep #3432 open until deployment, actual clean build, and rerun evidence are complete.