Skip to main content

Runbook — deploying migration 224 (chain_seq) without paging a false tamper

Target of the runbook_url on AuditChainBroken and AuditChainRolloutSkew in deploy/alerts/audit-chain.yaml.

Owner: platform / devops. Issues: #2968 (this runbook) · #2949 (the column) · #2963 (its review) · #2915 (the alerts).


TL;DR for a deployer​

Deploy migration 224 and the binaries that write audit rows together, in one change. Do not run the migration ahead of the rollout and let old replicas drain behind it.

If you do split them, expect AuditChainRolloutSkew (warning) for every affected org until the rollout completes and the next nightly verifier pass runs. That alert is not a security incident. AuditChainBroken (critical) still is.


Why the order matters​

Migration 224 added agent_audit_log.chain_seq and made it — not created_at — the key the audit hash chain is ordered on. created_at is stamped by the caller, before it contends for the append path's advisory lock, so under a burst rows commit in lock order while carrying contention-order timestamps and the chain forks. chain_seq is assigned inside the lock and is monotone in commit order by construction. (#2949; the measurement is in the migration header — 40 concurrent appends produced 16 distinct prev_hash values where 1 was correct.)

The column is NOT NULL DEFAULT 0, and the migration backfills every existing chained row with its position in the current walk order. So after the migration:

writerchain_seq it writeseffect
post-224 binarytail + 1, inside the lockcorrect
pre-224 binary0, the column defaultsorts before every backfilled row

Zero sorts first. The walk therefore starts at a row whose prev_hash is not the 32-zero chain root, invariant 1 fails, and the chain reports as broken.

The break is real and the loud posture is correct — for an audit chain a false alarm beats silent acceptance, and the window is bounded by the rollout. What #2968 fixed is that the alert used to describe it as tampering.


What the on-call sees now​

The verifier classifies each break. A link/order failure (invariants 1, 2 and 4) at a row carrying chain_seq = 0 is the rollout signature and nothing else can produce it; a content-hash failure (invariant 3) recomputes from the row's own columns, is independent of walk order, and is never a rollout artefact. That classification is exported as audit_chain_verifier_unsequenced_break, and the alert rules route on it:

you seemeansseverity
AuditChainRolloutSkewbreak at chain_seq = 0warning
AuditChainBrokenanything elsecritical

The two partition audit_chain_verifier_status == 0 per (org_id, epoch, chain_scope): no broken chain is silent, and no broken chain pages twice. A rollout on one org or one grain does not excuse a genuine break on another.

If the discriminator series is missing entirely — an old verifier binary, i.e. a mixed-version rollout of the thing doing the reporting — the join matches nothing and AuditChainBroken fires exactly as it did before #2968. Absence fails loud.


Deploying 224 correctly​

  1. Ship the migration and the writer binaries in one release. The writers of chained rows are the session-chain path and the org-grain path, both in internal/runtime/audit/pgstore.go. Anything that calls audit.EnableChain(...) is a writer.
  2. Roll the writers before or with the migration, never long after. A post-224 binary running against a pre-224 schema is safe — the column does not exist, and the code paths that use it are the ones the migration adds. The unsafe direction is the other one.
  3. If a canary or a partial rollout is unavoidable, say so in the deploy channel before starting, so the resulting AuditChainRolloutSkew is expected rather than investigated.
  4. The window closes when every writing replica is post-224 and the next nightly pass has re-walked the chains. The verifier runs at 02:00 UTC (deploy/cron/audit-verify.yaml), and the gauge holds its last value for a full day, so the warning will not clear before then.

Triaging AuditChainRolloutSkew​

  1. Confirm a rollout is actually in flight.

    kubectl get pods -n platform -o \
    jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[*].image}{"\n"}{end}'

    Every replica that writes audit rows must be on a post-224 image.

  2. If no rollout is in flight, this is not skew. Nothing else legitimately writes chain_seq = 0 to a chain that already has sequenced rows. Treat it exactly as AuditChainBroken and follow that runbook, including the escalation.

    This matters because chain_seq is deliberately not part of the row hash (see the migration header for why), so someone who can UPDATE the table can zero a row's sequence and choose which alert they trip.

    The verifier checks the breaking row's own content hash before it will classify an unsequenced row as an artefact, so they cannot edit a row's content and zero that row's sequence to hide it. They can hide behind a different row, and it takes two statements and no hash recomputation:

    1. edit row K's content -> K now fails the content check
    2. zero row J's chain_seq, J after K -> J sorts to the front of the walk

    J is untouched and self-consistent, so it passes the guard, fires the first-row invariant at chain_seq = 0, and the walk stops there without ever reaching K. That is inherent to "the walk stops at the first failure" plus "chain_seq is not hashed"; it is a severity downgrade, never a silencing; and step 1 is what catches it. A skew alert with no rollout in flight is the only signal you get for this class — treat it accordingly.

    Pinned by TestAuditChain2968_ATamperCanHideBehindALaterZeroedSequence_KnownLimit, which will fail if the limit is ever closed.

  3. Read the evidence.

    kubectl logs -n platform job/audit-verify-<n> | jq 'select(.ok == false)'

    Each line carries broken_at_unsequenced_row, broken_row_id, broken_at_row and broken_reason.

  4. Complete the rollout and wait for the next pass. Do not re-run the 224 backfill to "repair" the rows: it is idempotent and numbers only rows that need it, but re-running it while a pre-224 replica is still writing simply re-creates the condition one row later.


What NOT to do​

  • Do not silence either alert to make a dashboard green. AuditChainBroken was named as a live control in six code comments and existed in zero rules files for four years (#2873). Silencing restores that state.
  • Do not relax the AuditChainBroken predicate because it is noisy. If it is firing on something that is not tampering, the fix is a discriminator like this one, computed at the source — not a wider threshold.
  • Do not re-chain, truncate or "repair" a chain that has broken for any reason other than confirmed rollout skew. The broken link is the evidence.