Runbook — deploying migration 224 (chain_seq) without paging a false tamper
Target of the runbook_url on AuditChainBroken and AuditChainRolloutSkew in
deploy/alerts/audit-chain.yaml.
Owner: platform / devops. Issues: #2968 (this runbook) · #2949 (the column) · #2963 (its review) · #2915 (the alerts).
TL;DR for a deployer
Deploy migration 224 and the binaries that write audit rows together, in one change. Do not run the migration ahead of the rollout and let old replicas drain behind it.
If you do split them, expect AuditChainRolloutSkew (warning) for every affected
org until the rollout completes and the next nightly verifier pass runs. That
alert is not a security incident. AuditChainBroken (critical) still is.
Why the order matters
Migration 224 added agent_audit_log.chain_seq and made it — not created_at —
the key the audit hash chain is ordered on. created_at is stamped by the
caller, before it contends for the append path's advisory lock, so under a
burst rows commit in lock order while carrying contention-order timestamps and
the chain forks. chain_seq is assigned inside the lock and is monotone in
commit order by construction. (#2949; the measurement is in the migration
header — 40 concurrent appends produced 16 distinct prev_hash values where 1
was correct.)
The column is NOT NULL DEFAULT 0, and the migration backfills every existing
chained row with its position in the current walk order. So after the migration:
| writer | chain_seq it writes | effect |
|---|---|---|
| post-224 binary | tail + 1, inside the lock | correct |
| pre-224 binary | 0, the column default | sorts before every backfilled row |
Zero sorts first. The walk therefore starts at a row whose prev_hash is not
the 32-zero chain root, invariant 1 fails, and the chain reports as broken.
The break is real and the loud posture is correct — for an audit chain a false alarm beats silent acceptance, and the window is bounded by the rollout. What #2968 fixed is that the alert used to describe it as tampering.
What the on-call sees now
The verifier classifies each break. A link/order failure (invariants 1, 2
and 4) at a row carrying chain_seq = 0 is the rollout signature and nothing
else can produce it; a content-hash failure (invariant 3) recomputes from the
row's own columns, is independent of walk order, and is never a rollout
artefact. That classification is exported as
audit_chain_verifier_unsequenced_break, and the alert rules route on it:
| you see | means | severity |
|---|---|---|
AuditChainRolloutSkew | break at chain_seq = 0 | warning |
AuditChainBroken | anything else | critical |
The two partition audit_chain_verifier_status == 0 per
(org_id, epoch, chain_scope): no broken chain is silent, and no broken chain
pages twice. A rollout on one org or one grain does not excuse a genuine
break on another.
If the discriminator series is missing entirely — an old verifier binary, i.e.
a mixed-version rollout of the thing doing the reporting — the join matches
nothing and AuditChainBroken fires exactly as it did before #2968. Absence
fails loud.
Deploying 224 correctly
- Ship the migration and the writer binaries in one release. The writers of
chained rows are the session-chain path and the org-grain path, both in
internal/runtime/audit/pgstore.go. Anything that callsaudit.EnableChain(...)is a writer. - Roll the writers before or with the migration, never long after. A post-224 binary running against a pre-224 schema is safe — the column does not exist, and the code paths that use it are the ones the migration adds. The unsafe direction is the other one.
- If a canary or a partial rollout is unavoidable, say so in the deploy channel
before starting, so the resulting
AuditChainRolloutSkewis expected rather than investigated. - The window closes when every writing replica is post-224 and the next
nightly pass has re-walked the chains. The verifier runs at 02:00 UTC
(
deploy/cron/audit-verify.yaml), and the gauge holds its last value for a full day, so the warning will not clear before then.
Triaging AuditChainRolloutSkew
-
Confirm a rollout is actually in flight.
kubectl get pods -n platform -o \jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[*].image}{"\n"}{end}'Every replica that writes audit rows must be on a post-224 image.
-
If no rollout is in flight, this is not skew. Nothing else legitimately writes
chain_seq = 0to a chain that already has sequenced rows. Treat it exactly asAuditChainBrokenand follow that runbook, including the escalation.This matters because
chain_seqis deliberately not part of the row hash (see the migration header for why), so someone who canUPDATEthe table can zero a row's sequence and choose which alert they trip.The verifier checks the breaking row's own content hash before it will classify an unsequenced row as an artefact, so they cannot edit a row's content and zero that row's sequence to hide it. They can hide behind a different row, and it takes two statements and no hash recomputation:
1. edit row K's content -> K now fails the content check2. zero row J's chain_seq, J after K -> J sorts to the front of the walkJ is untouched and self-consistent, so it passes the guard, fires the first-row invariant at
chain_seq = 0, and the walk stops there without ever reaching K. That is inherent to "the walk stops at the first failure" plus "chain_seqis not hashed"; it is a severity downgrade, never a silencing; and step 1 is what catches it. A skew alert with no rollout in flight is the only signal you get for this class — treat it accordingly.Pinned by
TestAuditChain2968_ATamperCanHideBehindALaterZeroedSequence_KnownLimit, which will fail if the limit is ever closed. -
Read the evidence.
kubectl logs -n platform job/audit-verify-<n> | jq 'select(.ok == false)'Each line carries
broken_at_unsequenced_row,broken_row_id,broken_at_rowandbroken_reason. -
Complete the rollout and wait for the next pass. Do not re-run the 224 backfill to "repair" the rows: it is idempotent and numbers only rows that need it, but re-running it while a pre-224 replica is still writing simply re-creates the condition one row later.
What NOT to do
- Do not silence either alert to make a dashboard green.
AuditChainBrokenwas named as a live control in six code comments and existed in zero rules files for four years (#2873). Silencing restores that state. - Do not relax the
AuditChainBrokenpredicate because it is noisy. If it is firing on something that is not tampering, the fix is a discriminator like this one, computed at the source — not a wider threshold. - Do not re-chain, truncate or "repair" a chain that has broken for any reason other than confirmed rollout skew. The broken link is the evidence.
Related
deploy/alerts/audit-chain.yaml— the five rules and the reasoning behind each predicate.internal/context/store/migrations/224_audit_log_chain_seq.up.sql— the column, the indexes and the backfill.test/integration/context/audit_chain_2968_test.go— the classification, proved both ways against a real database.deploy/alerts/tests/audit-chain_test.yamlT10–T13 — the routing, proved both ways in promtool.