# wf-reconcile-drift.yaml — Prometheus alerting rules for the WF-31 workflow
# drift reconciler (internal/workflow/wfreconcile, run by cmd/reconciler).
#
# ── WHY THIS FILE EXISTS (issue #2495, architect finding F1 on PR #2488) ─────
#
# `wf_reconcile_drift_detected_total` is incremented per drift class and NOTHING
# WATCHED IT — at any class, since the reconciler shipped. A reconciler that
# detects drift and tells nobody is a detection control in name only, and the
# whole premise of WF-31 is that drift is RARE, which is precisely the condition
# under which nobody thinks to open the dashboard.
#
# The `orphaned_approval` class (#2381, added in PR #2488) makes it concrete: it
# detects `pending` governance approvals belonging to workflow runs that no
# longer exist. An approver sees a live, clearance-gated card for work that does
# not exist, answers it in good faith, and the decision lands in the ledger
# against a dead run — an audit-coherence breach against the "every action is
# auditable" tenet. If the #2381 interpreter fix ever regresses, this counter is
# what notices.
#
# ── THE COUNTER WAS ALSO NOT BEING EXPORTED. FIXED IN THE SAME PR. ──────────
#
# Do not skip this paragraph; without it every rule below is shelfware.
#
# `wfreconcile` instruments via `otel.Meter(...)`, which resolves against the
# GLOBAL OTel MeterProvider. `cmd/reconciler` never installed one, so the default
# NO-OP provider swallowed every wf_reconcile_* sample and the series could not
# appear on /metrics at all — regardless of any collector downstream. (Its
# sibling, internal/orgunit/dualwrite, uses promauto and so was always scrapable;
# that asymmetry is why deploy/alerts/dualwrite.yaml worked and this did not.)
# It was worse than "not scraped": the base binary did not LINK
# internal/runtime/metrics, go.opentelemetry.io/otel/sdk/metric or the Prometheus
# exporter at all, so the series was unobservable at link level, not merely
# unexported.
#
# cmd/reconciler/main.go now installs the provider. TWO guards pin it, and they
# pin DIFFERENT things — read both names before citing either:
#
#   cmd/reconciler/metrics_export_test.go
#     pins the MECHANISM. Negative control: the series is asserted ABSENT under
#     the default no-op provider and PRESENT after one is installed. It
#     constructs the provider ITSELF, so it says nothing about this binary's
#     wiring — delete the call from run() and it stays GREEN (measured).
#
#   cmd/reconciler/meterprovider_callsite_test.go
#     pins the CALL SITE, and is the one A4 below refers to. A static AST oracle:
#     run() must contain a call to <import>.NewProvider, resolved through the
#     import's local name so renaming the alias is a refactor and not a red.
#     A grep would be satisfied by the identifier appearing in a comment, in
#     dead code or in a test; the parser is not.
#
# The exact exposed series, MEASURED by that test rather than assumed (the OTel
# Prometheus exporter rewrites instrument names, so this is not a safe guess):
#
#   wf_reconcile_drift_detected_total{class="orphaned_approval",
#     otel_scope_name="context-engine/workflow",otel_scope_schema_url="",
#     otel_scope_version=""} 3
#
# Those `otel_scope_*` labels are why every rule below aggregates with
# `sum by (class)` BEFORE any binary operation. A raw `detected - converged`
# would join on the full label set; today the scope labels happen to be equal on
# both sides, so it would work by luck and break silently the day either
# instrument moved scope. That is the #2307 A1 defect exactly — an `and`/`-` whose
# label sets never match returns empty, and an empty rule is a dead rule with no
# error and green CI.
#
# ── SHELFWARE GATE: every metric referenced is emitted by real production code ─
#
#   wf_reconcile_drift_detected_total   -> internal/workflow/wfreconcile/metrics.go
#                                          (metrics.emitClass, per class per sweep)
#   wf_reconcile_drift_converged_total  -> same file, same function
#   wf_reconcile_sweeps_total           -> same file (metrics.observeSweep)
#
# `emitClass` SKIPS zero counts, so an idle class emits NO series at all. That is
# load-bearing for these rules in a good way: a class only becomes visible once
# it has actually detected something, and `increase()` returns to 0 within the
# window once detection stops — so every alert here self-resolves without a
# manual reset.
#
# ── MIRRORED TO dev/prometheus/rules/ SINCE #3193 — AND A4 FIRES THERE ───────
#
# This header used to say the opposite, and the reason it gave was sound in
# isolation: cmd/reconciler is NOT a service in docker-compose.dev.yml (it is
# published to GHCR and runs on GKE), so mirroring arms A4's `absent()` arm
# permanently, and noise is how a real alert gets ignored.
#
# What that reasoning could not see is what the exclusion cost. DEV_MIRROR was
# an opt-in allow-list AND `--check` iterated the allow-list rather than the
# CRDs on disk, so 4 of 5 CRDs — this one included — were evaluated by no
# running Prometheus while CI printed "dev rule mirrors are in sync" and exited
# 0 (#3193). Absence scored as a pass. The exclusion was invisible, and an
# invisible exclusion is indistinguishable from an oversight.
#
# So: mirrored. A4 fires on beta-dev, permanently, and the statement it makes
# there — "the WF-31 drift reconciler is not running against this database" — is
# TRUE. There is no Alertmanager in docker-compose.dev.yml, so that costs a red
# row on the /alerts page and nobody's sleep. IF AN ALERTMANAGER IS EVER
# ATTACHED TO THE COMPOSE STACK, route or silence A4 in the same change, or move
# this file to NOT_MIRRORED in scripts/render-alert-rules.py with that as the
# recorded reason.
#
# These rules are linted and UNIT-TESTED in CI either way: check-alert-rules.sh
# renders every CRD and deploy/alerts/tests/wf-reconcile-drift_test.yaml runs
# against the render.
#
# References:
#   - issue #2495 (this file), architect review 4862636135 on PR #2488
#   - issue #2381 (the orphaned-approval leak), #2488 (the drift class)
#   - issue #2583 (an observed `applyToSession … session not found` on beta-dev,
#     evidence since lost) — A1 is the alarm that would have made that state
#     observable rather than a log line nobody tailed.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: wf-reconcile-drift
  namespace: platform
  labels:
    app: reconciler
    role: alert-rules
    release: kube-prometheus-stack
spec:
  groups:
    # ──────────────────────────────────────────────────────────────
    # Recording rules.
    #
    # Every alert below reads one of these two series rather than
    # re-deriving, so the aggregation (and therefore the label set)
    # is written ONCE. See the header for why that matters.
    # ──────────────────────────────────────────────────────────────
    - name: wf-reconcile.recording
      interval: 5m
      rules:
        # Drift items detected per class over a window that is
        # guaranteed to contain at least one sweep.
        #
        # 26h, not 24h. The reconcile cadence is
        # wfreconcile.DefaultInterval = 24h with an IMMEDIATE first
        # sweep, so a 24h window can align exactly with the sweep
        # period and — with rate/increase extrapolation at the window
        # edges — straddle the boundary such that a real increment
        # lands outside it. The 2h margin is the same correction
        # memory-embed's A7 documents for its `offset 26h`.
        - record: wf_reconcile_drift_detected_sweep
          expr: sum by (class) (increase(wf_reconcile_drift_detected_total[26h]))

        # Drift detected but NOT converged, per class, over the same
        # window.
        #
        # The `or … * 0` is the fill-missing-with-zero idiom and it is
        # required, not decorative: a class that has detected drift but
        # never converged any emits NO `converged` series at all
        # (emitClass skips zeros), and `A - B` over a missing B returns
        # EMPTY, not A. Without the fill this rule would be silent in
        # exactly the situation it exists to catch. Pinned by T3 in the
        # fixture.
        - record: wf_reconcile_drift_unconverged_sweep
          expr: |
            wf_reconcile_drift_detected_sweep
            -
            (
              sum by (class) (increase(wf_reconcile_drift_converged_total[26h]))
              or
              wf_reconcile_drift_detected_sweep * 0
            )

    # ──────────────────────────────────────────────────────────────
    # Alerting rules.
    # ──────────────────────────────────────────────────────────────
    - name: wf-reconcile.alerts
      rules:
        # ──────────────────────────────────────────────────────────
        # A1. ANY orphaned pending approval. Non-zero is the alarm.
        #
        # This class is DETECT-ONLY by design (approvalreconcile.go:
        # voiding a clearance-gated governance row from an unattended
        # sweeper is a decision for an ADR and a founder, not a bug-fix
        # PR). So `Converged` is permanently 0 for it and A2 below must
        # exclude it — otherwise A2 would fire forever and mean nothing.
        # A1 is that class's alarm instead.
        #
        # WARNING, not CRITICAL, and the reasoning is recorded so it can
        # be argued with: the breached state is DURABLE, not
        # time-critical. Paging at 03:00 cannot make an already-orphaned
        # approval less orphaned, and the remediation — reading the ids
        # this class logs and voiding them through the operator
        # VoidApproval API — is a daytime action with a human principal.
        # What it must never be is SILENT, which is what it was. Route
        # this to a ticket queue, not to the pager.
        #
        # `for: 10m` is debounce only. The counter jumps once per sweep
        # and the 26h window then holds the value up for ~26h, so this
        # does not delay detection meaningfully; it just prevents a
        # single missed scrape from flapping the alert.
        # ──────────────────────────────────────────────────────────
        - alert: WorkflowReconcileOrphanedApprovalDetected
          expr: wf_reconcile_drift_detected_sweep{class="orphaned_approval"} > 0
          for: 10m
          labels:
            severity: warning
            component: wf-reconciler
            team: platform
            governance: "true"
          annotations:
            summary: "{{ $value }} orphaned pending approval(s) detected by the WF-31 reconciler"
            description: |
              The WF-31 reconciler found `pending` rows in
              `governance_approvals` whose owning workflow run is
              TERMINAL or GONE (#2381). Each one renders in the
              Approvals queue as a live, clearance-gated card for work
              that no longer exists. An approver can answer it in good
              faith and their decision lands in the ledger against a
              dead run — an audit-coherence breach, not a cosmetic one.

              This class converges NOTHING by design, so this alert will
              not clear on its own until the rows are dealt with and a
              subsequent sweep finds none.

              What to do:
                1. Read the reconciler's WARN lines. Each carries
                   `org_id`, `approval_id`, `action_type`, `run_id` and
                   `run_status` (`<run row deleted>` when the run is
                   gone, which is the worse of the two breaches — no
                   subsequent event can ever resolve that row).
                2. Void each id through the operator VoidApproval API,
                   which is audited with a real principal.
                3. If the count is GROWING sweep over sweep rather than
                   draining, the #2381 interpreter fix has regressed and
                   NEW leaks are being created. That is a P1, not a
                   cleanup.
            runbook_url: https://github.com/upsquad-ai/upsquad-core/blob/main/docs/runbooks/wf-reconcile-drift.md

        # ──────────────────────────────────────────────────────────
        # A2. Drift detected that the reconciler could NOT converge.
        #
        # Every class except orphaned_approval both detects AND heals,
        # within the same sweep. Detected-minus-converged should be 0.
        # A non-zero difference means the heal path is erroring — the
        # reconciler is finding the problem and failing to fix it, which
        # is a strictly worse state than not looking, because the
        # dashboard says the reconciler ran.
        #
        # `class!="orphaned_approval"` is LOAD-BEARING, not tidiness:
        # that class is detect-only, so including it would pin this
        # alert permanently on and destroy its meaning.
        # ──────────────────────────────────────────────────────────
        - alert: WorkflowReconcileDriftNotConverging
          expr: wf_reconcile_drift_unconverged_sweep{class!="orphaned_approval"} > 0
          for: 1h
          labels:
            severity: warning
            component: wf-reconciler
            team: platform
          annotations:
            summary: "WF-31 drift class {{ $labels.class }} detected {{ $value }} item(s) it could not converge"
            description: |
              `{{ $labels.class }}` detected drift and did not heal all
              of it in the same sweep. Convergence failures also bump
              `workflow_schedule_reconcile_failures_total` (shared with
              the WF-28 post-commit hook), so check that counter and the
              reconciler's error log together.

              Per-class starting points:
                missing_schedule / paused_mismatch  — the Temporal
                  schedule reissue is failing. Check Temporal
                  reachability and the namespace.
                orphaned_schedule — the prune is failing. NOTE that
                  Temporal's schedule LIST is a visibility index updated
                  asynchronously after a delete, so a single sweep can
                  legitimately re-report a schedule it just pruned; the
                  1h `for:` is sized to absorb that.
                stale_projection — the run-status heal is failing.
                orphaned_task / stranded_task — the coordinator task
                  terminalizer is refusing the transition; the graph
                  guard is fail-closed by design, so read the audit hop
                  before assuming the reconciler is at fault.
            runbook_url: https://github.com/upsquad-ai/upsquad-core/blob/main/docs/runbooks/wf-reconcile-drift.md

        # ──────────────────────────────────────────────────────────
        # A3. Drift keeps coming back, sweep after sweep.
        #
        # Distinct from A2, and both are wanted. A2 is "the reconciler
        # cannot fix it". A3 is "the reconciler fixes it and something
        # upstream keeps re-creating it" — convergence hides the
        # underlying defect, so the only evidence is that the sweeper
        # keeps having work to do. The reconciler is a safety net, and a
        # safety net catching something every single day is a bug
        # somewhere else.
        #
        # `offset 2d` rather than min_over_time/count_over_time: it
        # STRUCTURALLY requires two days of history. With no sample at
        # t-2d the right-hand selector is empty, the `and` yields empty,
        # and the alert cannot fire — so a series that has only just
        # appeared cannot false-positive its way into a "sustained"
        # alert on its first day. Fails silent, never loud, which is the
        # correct direction for a warn-level churn signal. Pinned by T5
        # and T6.
        # ──────────────────────────────────────────────────────────
        - alert: WorkflowReconcileDriftSustained
          expr: |
            (wf_reconcile_drift_detected_sweep{class!="orphaned_approval"} > 0)
            and
            (wf_reconcile_drift_detected_sweep{class!="orphaned_approval"} offset 2d > 0)
          for: 30m
          labels:
            severity: warning
            component: wf-reconciler
            team: platform
          annotations:
            summary: "WF-31 drift class {{ $labels.class }} has been non-zero for 2+ days"
            description: |
              `{{ $labels.class }}` has detected drift both now and two
              days ago. The reconciler may well be converging it every
              time — that is what makes this quiet and what makes it
              worth an alert. Sustained drift on a class that heals
              means the write path that is SUPPOSED to keep these in
              step is failing and the sweeper is papering over it.

              Find the producer, not the sweeper. Check the WF-28
              post-commit schedule hook and the interpreter's run-status
              projection before touching anything in
              internal/workflow/wfreconcile.
            runbook_url: https://github.com/upsquad-ai/upsquad-core/blob/main/docs/runbooks/wf-reconcile-drift.md

        # ──────────────────────────────────────────────────────────
        # A4. The reconciler is not sweeping at all.
        #
        # THE MOST IMPORTANT RULE IN THIS FILE, and the one that is easy
        # to leave out. A1-A3 all silently assume the sweeper runs. A
        # reconciler that has stopped emits NO drift metrics, which is
        # indistinguishable from a perfectly healthy system — the exact
        # failure mode #2315 hit when the beta-dev Prometheus was found
        # evaluating ZERO rules off a stale mount and everything looked
        # fine.
        #
        # Two arms because one is not enough:
        #   * `increase(...[50h]) == 0` — the process is alive and
        #     scraping but the loop is wedged (deadlock, ctx cancelled,
        #     advisory lock held elsewhere). 50h ~= two 24h sweeps.
        #   * `absent(...)` — the series is gone entirely: the pod is
        #     down, crash-looping, or the scrape target vanished. The
        #     first arm CANNOT see this, because a stale series leaves
        #     the instant vector empty and an empty vector fires nothing.
        #
        # `absent()` also fires where the reconciler is deliberately not
        # deployed — which is the case on the beta-dev compose stack, so
        # this arm is permanently firing there since #3193 mirrored the
        # file. That is intended: "the drift detector is switched off" is
        # a state that should be loud, and on beta-dev it is true. See the
        # file header for what to do if an Alertmanager is ever attached.
        #
        # The first sweep runs IMMEDIATELY on boot (RunLoop, not after
        # one interval), so there is no 24h startup blind window to
        # absorb; `for: 1h` covers a rollout.
        # ──────────────────────────────────────────────────────────
        - alert: WorkflowReconcileSweepsStalled
          expr: |
            (increase(wf_reconcile_sweeps_total[50h]) == 0)
            or
            absent(wf_reconcile_sweeps_total)
          for: 1h
          labels:
            severity: critical
            component: wf-reconciler
            team: platform
          annotations:
            summary: "The WF-31 drift reconciler has not completed a sweep in 2+ days"
            description: |
              `wf_reconcile_sweeps_total` has not advanced in over two
              sweep intervals, or the series is absent altogether.

              While this is firing, treat A1-A3 as MEANINGLESS rather
              than as "no drift": a sweeper that is not running reports
              exactly the same thing as a healthy system.

              Check, in order:
                1. `kubectl -n platform logs deploy/reconciler` — the
                   daemon exits rather than waits if another replica
                   holds the `upsquad.orgunit_reconciler` advisory lock.
                2. `WF_RECONCILE_DISABLED` — the kill-switch for this
                   whole surface. It logs a WARN on startup when set.
                3. That the OTel MeterProvider is still wired in
                   cmd/reconciler/main.go (#2495). Removing it makes
                   every wf_reconcile_* series vanish while the daemon
                   runs perfectly — which is what this alert would then
                   be reporting, and the wrong thing to chase.
                   cmd/reconciler/meterprovider_callsite_test.go is the
                   guard for THAT (a static AST oracle over run()).
                   metrics_export_test.go is NOT — it builds its own
                   provider and stays green when the call site is
                   deleted.
            runbook_url: https://github.com/upsquad-ai/upsquad-core/blob/main/docs/runbooks/wf-reconcile-drift.md
