# Audit hash-chain tamper-evidence alerting.
#
# Issues: #2873 (the rule did not exist) · #2316 (nothing loaded this directory)
# Task:   #2915 — MG slice-3 T13 · LLD #2901 §5 · §8 T13
# Source of the metrics: cmd/audit-verify, instruments in
#                        internal/runtime/metrics/instruments.go
#
# ── WHY THIS FILE IS NEW RATHER THAN A LINE ADDED SOMEWHERE ──────────────────
#
# `AuditChainBroken` was named as a live control in six code comments and twice
# in docs/lld/wave1-item4-audit-hash-chain.md, and was defined in none of the
# eight rules files in this repo. Six agreeing comments are ONE source: they all
# descend from the LLD line. The control was a phantom, and the tamper-evidence
# property the platform asserts to tenants ("every action is auditable, the
# chain is verified nightly") rested on it.
#
# Two further layers of the same defect were found writing this file, and both
# are closed here rather than left for the reader to trip over:
#
#   1. deploy/alerts/ was referenced by NO kustomization (#2316), so a rule
#      written here would have been documentation wearing the shape of
#      configuration. deploy/alerts/kustomization.yaml + the ArgoCD Application
#      in deployments/argocd/observability-dev.yaml fix that, and
#      scripts/check-rules-loaded.py fails CI if it regresses. The check asserts
#      the rule appears in a `kustomize build` of a real Application path — it
#      does not check that the file exists, because file-exists is exactly the
#      predicate that was already true while nothing was deployed.
#
#   2. cmd/audit-verify installed no OTel MeterProvider, so every Record() went
#      to the global no-op and the gauge did not exist as a series at all. Fixed
#      in the same change (cmd/audit-verify/main.go). An alert on a gauge
#      nothing writes is worse than no alert, because it reads as coverage.
#
# ── THE PREDICATE, AND WHY IT IS BARE `== 0` OVER THE RAW SERIES ─────────────
#
# The gauge emits TWO SERIES PER ORG since #2814/#2859:
#
#     audit_chain_verifier_status{org_id="…", epoch="1"}                       # session grain
#     audit_chain_verifier_status{org_id="…", epoch="1", chain_scope="model_gateway"}
#
# `chain_scope` is omitted, not empty, on the session grain — see
# cmd/audit-verify/metrics_attrs.go for why that is load-bearing.
#
# Three candidate shapes, and the two that were rejected:
#
#   * `min(audit_chain_verifier_status) == 0` — what Grafana panel 22 did. It
#     DOES detect (a global min goes to 0 if any series is 0), but it produces
#     ONE alert instance carrying NO labels, so the page cannot name the org or
#     the grain and the on-call has nowhere to start. ADR-0034 condition 2 names
#     this over-aggregation explicitly. Panel 22 is fixed in the same change.
#   * `min by (org_id) (...) == 0` — names the org, still collapses the grains,
#     so it sends an operator to the wrong service: a broken `model_gateway`
#     chain and a broken session chain are owned by different code.
#   * `audit_chain_verifier_status == 0` — one alert per series, every label
#     preserved, and a THIRD grain added later joins by existing rather than by
#     someone remembering to widen a `by` clause. This one.
#
# An instant-vector selector matches on a label SUBSET, so this expression
# covers both series today and any future grain without amendment. That is the
# property #2859's label fix exists to create, and this is the consumer that
# makes it worth having.
#
# ── "STOPPED RUNNING" IS NOT "FOUND NOTHING" ────────────────────────────────
#
# `AuditChainBroken` can only fire on a value the verifier WROTE. Its silence is
# therefore ambiguous, and the ambiguity resolves the dangerous way by default:
# a suspended CronJob, a bad image tag, a failed pod, or a dead metrics pipeline
# all look exactly like a clean bill of health. Three alerts, three states, and
# the discriminator is a liveness gauge the status gauge cannot express:
#
#   AuditChainBroken                 status == 0, not skew   ran, found corruption
#   AuditChainRolloutSkew            unsequenced_break == 1  ran, found a 224 deploy
#   AuditChainVerifierStale          last_run older than 26h ran once, has stopped
#   AuditChainVerifierMetricsAbsent  absent(last_run)        never arrived at all
#   AuditChainVerifierWalkErrors     errors_total > 0        ran, could not look
#
# The first two PARTITION `status == 0` — see the #2968 section below. Together
# they cover it exactly, so adding the second one took nothing away from the
# first.
#
# The absent() arm is what makes this ruleset fail LOUD instead of fail SILENT.
# Either the verifier's metrics arrive (and the `== 0` arm is live) or they do
# not (and the absent arm pages). There is no configuration of the world in
# which the whole ruleset is quietly inert — which is the exact state #2873 and
# #2316 found it in, one and two levels down respectively.
#
# ── ON beta-dev, THAT ARM IS PERMANENTLY FIRING — DELIBERATELY (#3193) ───────
#
# This CRD is mirrored to dev/prometheus/rules/audit-chain.yml and loaded by the
# compose Prometheus behind beta-dev. No cmd/audit-verify CronJob runs against
# that database, so `AuditChainVerifierMetricsAbsent` matches there and stays
# matched. That is the paragraph above being right, not a defect: a stack where
# nobody verifies the audit chain should say so out loud, and on beta-dev nobody
# does. Measured 2026-09-07 — it is the only arm of this ruleset that matches
# there; every other rule reads a series that does not exist and evaluates
# empty.
#
# There is no Alertmanager in docker-compose.dev.yml, so today that costs a red
# row on the /alerts page and nobody's sleep. IF ONE IS EVER ATTACHED TO THE
# COMPOSE STACK, route or silence this arm in the same change — or move this
# file to NOT_MIRRORED in scripts/render-alert-rules.py with that as the
# recorded reason. wf-reconcile-drift.yaml carries the identical note for
# `WorkflowReconcileSweepsStalled`, the other arm in the same position.
#
# ── A DEPLOY WINDOW IS NOT A BREACH, AND THE ALERT NOW SAYS WHICH (#2968) ───
#
# Migration 224 added `agent_audit_log.chain_seq` and made it, not `created_at`,
# the key the chain is ordered on (#2949). A PRE-224 BINARY WRITING A CHAINED
# ROW AFTER THAT MIGRATION HAS RUN leaves the column at its 0 default. Zero
# sorts before every backfilled row, so the walk starts where the hashes do not
# and invariant 1 fails at the very first row. Every mixed-version rollout of
# that migration therefore produces a break — correctly, and the loud posture is
# the right one for an audit chain.
#
# What was wrong was the WORDING. `AuditChainBroken` says "no longer
# tamper-evident … escalate to the founder, a confirmed break is a disclosure
# question". Paging that at 02:00 for a deploy the operator started themselves
# is how a real detection gets distrusted, and the on-call had no way to tell
# the two apart from the alert alone (#2968; the mitigation existed only in the
# migration's file header and PR #2963's body, neither of which is an operator
# surface).
#
# THE DISCRIMINATOR IS COMPUTED AT THE SOURCE, NOT GUESSED IN PromQL.
# cmd/audit-verify emits `audit_chain_verifier_unsequenced_break` = 1 when the
# break it found was a LINK/ORDER failure (invariant 1, 2 or 4) at a row
# carrying chain_seq = 0, and 0 otherwise. Invariant 3 — the content hash — is
# deliberately excluded: it recomputes from the row's OWN prev_hash column and
# is independent of walk order, so a pre-224 row passes it and a row that fails
# it has had its content changed no matter where its chain_seq sits.
#
# The routing that buys, and the three properties that make it safe:
#
#   1. A REAL BREAK DURING A DEPLOY WINDOW STILL PAGES CRITICAL. The join is
#      per-series on (org_id, epoch, chain_scope), so skew on org A's session
#      chain does not touch org B's, and skew on one grain does not touch the
#      other. Fixture T10 is exactly this case.
#   2. NOTHING IS EVER SILENCED. The skew world fires AuditChainRolloutSkew
#      instead — warning, different wording, and it states outright that the
#      chain is UNVERIFIED rather than verified. `unless` moves a page between
#      two rules; it never removes the last one.
#   3. ABSENCE FAILS LOUD. If the discriminator series does not exist — an old
#      verifier binary, i.e. a mixed-version rollout of the verifier ITSELF —
#      `unless` matches nothing and AuditChainBroken fires exactly as it did
#      before this gauge was invented. Fixture T13.
#
# The residual, stated rather than hidden, and BROADER than the first draft of
# this paragraph claimed (corrected on review, with a counterexample):
#
# An insider who can UPDATE the table can set chain_seq = 0 and downgrade a
# tamper page from critical to warning. The anti-downgrade guard stops them
# doing it to the row they edited — it checks the BREAKING row's own content
# hash — but not to the chain. Two statements, no hash recomputed:
#
#     1. edit row K's content              -> K now fails invariant 3
#     2. zero row J's chain_seq, J after K -> J sorts to the FRONT
#
# J is untouched and self-consistent, so it passes the guard; invariant 1 fires
# at chain_seq = 0; and the walk stops there without ever reaching K. Measured
# on a real database and pinned by
# TestAuditChain2968_ATamperCanHideBehindALaterZeroedSequence_KnownLimit.
#
# Accepted rather than closed: it is inherent to "the walk stops at the first
# failure" plus "chain_seq is not hashed", it needs UPDATE on the audit table,
# and it is a severity DOWNGRADE — never a silencing. The compensating control
# is on the surface that matters: AuditChainRolloutSkew states outright that the
# chain is UNVERIFIED and that a break further along it is invisible, and its
# runbook step 2 sends an unexplained skew straight to the tamper path.
#
# Deploy order — migration and binary together — is recorded for deployers in
# docs/runbooks/audit-chain-224-rollout.md.
#
# ── RECORDED, NOT ABSORBED: #2005 (P0) INTERACTS WITH EVERY RULE HERE ────────
#
# The retention sweeper's REDACT path pseudonymises an org's entire audit set on
# every tick, hash-chained tables included. `AuditChainBroken` would therefore
# fire CORRECTLY — and constantly — the moment that sweeper is enabled, because
# the chain really would be broken by it. The sweeper is off and must stay off
# (#2005's own instruction). Do not "fix" a storm of this alert by relaxing the
# predicate; the predicate would be right and the sweeper would be wrong.
#
# ── TRANSPORT, STATED BECAUSE IT IS NOT YET SOLVED ──────────────────────────
#
# cmd/audit-verify is a CronJob. Its pod exits in seconds, so a pull-based
# scrape will essentially never catch it, and the SDK's Prometheus exporter is a
# pull reader. The values are RECORDED correctly as of this change; getting them
# to Prometheus needs a push path (Pushgateway, an OTLP collector, or exporting
# from a long-lived process) and that decision is escalated rather than invented
# here. Until it lands, AuditChainVerifierMetricsAbsent is the arm that fires,
# and it fires for a true reason.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: audit-chain
  namespace: platform
  labels:
    app: audit-verify
    role: alert-rules
    # The kube-prometheus-stack Prometheus ruleSelector matches on this. Same
    # value as every other PrometheusRule in this repo — a rule with the wrong
    # release label is loaded by nothing and is the CRD-shaped version of the
    # bug this file exists to close.
    release: kube-prometheus-stack
spec:
  groups:
    - name: audit-chain.alerts
      # One evaluation per minute is ample for a signal written once a day, and
      # keeps the `for:` windows below meaningful in samples rather than in one
      # lonely scrape.
      interval: 1m
      rules:
        - alert: AuditChainBroken
          # NO `by`, NO `min`, NO `count`. See the header.
          #
          # The `unless` is the #2968 discriminator and it is a ROUTE, not a
          # suppression: the series it removes are picked up by
          # AuditChainRolloutSkew below, so no world exists in which a broken
          # chain produces no alert. Three things about its shape:
          #
          #   * `on (org_id, epoch, chain_scope)` and not `on (org_id)`. Per
          #     series, so a rollout artefact on one org/grain cannot excuse a
          #     genuine break on another (T10).
          #   * `chain_scope` is named even though the session grain OMITS it.
          #     Vector matching reads a missing label as the empty string, so
          #     both sides match at "" on the session grain and on the real
          #     value on the scope grain — one clause, both grains, and a third
          #     grain joins by existing.
          #   * `== 1` on the right-hand side, not a bare selector. The gauge is
          #     emitted for EVERY chain including intact ones, so a bare
          #     selector would match the 0s too and remove every series.
          expr: |
            (audit_chain_verifier_status == 0)
              unless on (org_id, epoch, chain_scope)
            (audit_chain_verifier_unsequenced_break == 1)
          # Long enough to survive a single bad scrape, short enough that the
          # LLD's "pages the on-call within 5 min" claim becomes true rather
          # than aspirational.
          for: 5m
          labels:
            severity: critical
            component: audit
            # Which grain, as a label rather than only as prose, so the routing
            # tree can send org-grain breakage to the owning service later
            # without rewriting the rule.
            #
            # if/else, NOT `{{ $labels.chain_scope | default "session" }}`.
            # #2873's draft used the latter; `default` is a Sprig function and
            # is NOT registered in Prometheus alert templates, so it would fail
            # to render — and a template that fails to render fails inside the
            # notification, where nobody is looking. Plain text/template only.
            grain: '{{ if $labels.chain_scope }}scope{{ else }}session{{ end }}'
          annotations:
            summary: >-
              Audit hash chain BROKEN for org {{ $labels.org_id }}
              ({{ if $labels.chain_scope }}scope {{ $labels.chain_scope }}{{ else }}session grain{{ end }})
            description: >-
              cmd/audit-verify walked the chained audit rows for org
              {{ $labels.org_id }} at epoch {{ $labels.epoch }} and found a
              link whose stored hash does not match its recomputed value.
              Rows in this window are no longer tamper-evident. This is an
              integrity finding, not an availability one: nothing is down.
              A MIXED-VERSION ROLLOUT OF MIGRATION 224 HAS ALREADY BEEN RULED
              OUT — that case fires AuditChainRolloutSkew instead, and this
              rule excludes it per-series (#2968).
            runbook: >-
              0. Rollout skew is already excluded, so do not spend the first
              ten minutes there. If AuditChainRolloutSkew is ALSO firing, it is
              for a different org, epoch or chain grain than this alert: the
              two are joined per series, not per org.
              1. Do NOT truncate, re-chain or "repair" anything before the
              evidence is captured — the broken link IS the evidence.
              2. `kubectl logs job/audit-verify-<n> -n platform` — the JSON
              line for this org carries broken_session_id, broken_at_row,
              broken_row_id and broken_reason.
              3. Confirm the retention sweeper (#2005) is still disabled. Its
              REDACT path pseudonymises hash-chained rows and would produce
              this alert legitimately across EVERY org at once. A single org
              is tampering or corruption; all orgs at once is the sweeper.
              4. Escalate to the founder. A confirmed break is a disclosure
              question, not only an engineering one.
            runbook_url: docs/runbooks/audit-chain-224-rollout.md
            dashboard: ICS Overview, panel 22 (Audit Chain Verifier Status)

        - alert: AuditChainRolloutSkew
          # The other half of the #2968 route. Every series AuditChainBroken's
          # `unless` removes lands here, so the two together cover
          # `audit_chain_verifier_status == 0` exactly — no gap, no overlap.
          expr: audit_chain_verifier_unsequenced_break == 1
          # Same 5m as AuditChainBroken. The two rules describe one observation
          # and must reach the on-call at the same moment; a longer `for` here
          # would open a window in which a broken chain has been routed away
          # from the critical alert and has not yet arrived at this one.
          for: 5m
          labels:
            # WARNING, AND THIS IS THE WHOLE POINT OF #2968 — but read the
            # description before treating it as ignorable. A deploy the
            # operator started themselves is not a security incident and must
            # not page as one; a chain that is UNVERIFIED still is not healthy.
            # Warning is the severity that says "triage this in hours, not
            # seconds"; it is not silence, and #2968 explicitly did not ask for
            # silence.
            severity: warning
            component: audit
            grain: '{{ if $labels.chain_scope }}scope{{ else }}session{{ end }}'
          annotations:
            summary: >-
              Audit chain for org {{ $labels.org_id }}
              ({{ if $labels.chain_scope }}scope {{ $labels.chain_scope }}{{ else }}session grain{{ end }})
              broke at an UNSEQUENCED row — migration 224 rollout skew, not tampering
            description: >-
              cmd/audit-verify could not walk this chain, and the row it failed
              at carries chain_seq = 0. That is the signature of a binary
              PREDATING migration 224 writing a chained row after that
              migration ran: the column takes its 0 default, 0 sorts before
              every backfilled row, and the walk starts where the hashes do
              not. It is a deploy artefact. It is NOT a tamper finding — a
              content tamper fails invariant 3, which recomputes from the
              row's own prev_hash column and cannot be produced this way.
              READ THE SECOND HALF. This chain is UNVERIFIED, not verified:
              the walk stops at the first failure, so a genuine break further
              along it is invisible for as long as this is firing. Treat the
              silence of AuditChainBroken for this org and grain as no
              evidence until this clears.
            runbook: >-
              1. Confirm a rollout is actually in flight —
              `kubectl get pods -n platform -l app=<writer> -o
              jsonpath='{.items[*].spec.containers[*].image}'`. Every replica
              that writes audit rows must be on a post-224 image.
              2. IF NO ROLLOUT IS IN FLIGHT, THIS IS NOT SKEW. Nothing else
              legitimately writes chain_seq = 0 to a chain that already has
              sequenced rows. Escalate it as you would AuditChainBroken; an
              insider who can UPDATE the table can set chain_seq = 0 to
              downgrade their own page to this one.
              3. Complete the rollout, then wait for the next nightly pass
              (02:00 UTC) — the gauge holds its last value for a full day and
              clears only when a pass re-walks the chain cleanly.
              4. `kubectl logs job/audit-verify-<n> -n platform`: the JSON line
              for this org carries broken_at_unsequenced_row, broken_row_id and
              broken_reason.
            runbook_url: docs/runbooks/audit-chain-224-rollout.md
            dashboard: ICS Overview, panel 22 (Audit Chain Verifier Status)

        - alert: AuditChainVerifierStale
          # 26h, not 24h: the job runs daily, so a 24h threshold would flap on
          # ordinary scheduling jitter and a 2h grace is cheaper than a rule
          # nobody trusts. If the schedule in deploy/cron/audit-verify.yaml
          # changes, this number is downstream of it.
          expr: time() - audit_chain_verifier_last_run_timestamp_seconds > 93600
          for: 30m
          labels:
            severity: critical
            component: audit
          annotations:
            summary: Audit hash-chain verifier has not completed a pass in over 26h
            description: >-
              The last completed verifier pass was
              {{ $value | humanizeDuration }} ago. AuditChainBroken cannot fire
              while this is true — it is keyed on a value only a completed pass
              writes — so the audit chain is currently UNVERIFIED, not verified
              and healthy. Treat the absence of AuditChainBroken as no evidence
              for as long as this alert is firing.
            runbook: >-
              `kubectl get cronjob audit-verify -n platform` (suspended?),
              then `kubectl get jobs -n platform -l app=audit-verify` for the
              last few runs. Exit 2 is an infrastructure/config error; exit 1
              is a genuine chain break that should also have paged.

        - alert: AuditChainVerifierMetricsAbsent
          # The fail-loud arm. `absent()` returns a series only when the
          # selector matches NOTHING, so this is the one alert in the file that
          # fires when the pipeline that feeds the other three is dead.
          #
          # 6h rather than minutes: a fresh Prometheus, a cluster rebuild, or a
          # deploy window must not page. Six hours is well inside the daily
          # cadence and well outside any restart.
          expr: absent(audit_chain_verifier_last_run_timestamp_seconds)
          for: 6h
          labels:
            severity: critical
            component: audit
          annotations:
            summary: Audit hash-chain verifier metrics are not reaching Prometheus at all
            description: >-
              No audit_chain_verifier_last_run_timestamp_seconds series exists.
              Either cmd/audit-verify has never run in this environment, or its
              metrics never reach Prometheus. cmd/audit-verify is a CronJob and
              its pod exits in seconds, so a pull-based scrape will not catch
              it; a push path is required and is tracked as a separate gate on
              the :8091 promotion. Until that lands this alert firing is the
              CORRECT reading of the world, and every other alert in this
              file is silent for a reason that is not health.
            runbook: >-
              Confirm the CronJob exists and has run, then confirm the metrics
              transport. Do not silence this alert to make a dashboard green —
              silencing it restores the exact state #2873 documented.

        - alert: AuditChainVerifierWalkErrors
          # Not `increase(...[26h]) > 0`: with one sample per day a 26h range
          # vector frequently holds a single point, and increase() over one
          # point is empty — the rule would be structurally incapable of firing
          # for the ordinary schedule. Every run rewrites this counter from a
          # fresh process, so the current value IS the last run's error count,
          # and cmd/audit-verify emits 0 explicitly on a clean walk so a stale
          # non-zero cannot stand.
          expr: audit_chain_verifier_errors_total > 0
          for: 30m
          labels:
            severity: warning
            component: audit
          annotations:
            summary: >-
              Audit chain verifier could not walk org {{ $labels.org_id }}
              ({{ $labels.stage }} grain)
            description: >-
              The walk failed for infrastructure reasons — a DB error, a
              timeout — so this org's chain was neither verified nor found
              broken. The status gauge is deliberately NOT flipped to 0 for
              this, because 0 means tampering and a DB timeout is not
              tampering. Warning rather than critical: nothing is known to be
              wrong, but nothing is known to be right either, and a
              persistently erroring walk is a silent blind spot.
            runbook: >-
              Job logs carry the underlying error per org. A single org
              erroring points at that org's data volume or a lock; all orgs
              erroring points at the DB connection or credentials.
