Skip to main content

Runbook — WF-31 workflow drift reconciler alerts

Target of the runbook_url on every alert in deploy/alerts/wf-reconcile-drift.yaml.

Owner: platform / devops. Issue: #2495.


What the reconciler is, and what it is not​

internal/workflow/wfreconcile, run by the cmd/reconciler daemon, sweeps every tenant once a day (WF_RECONCILE_INTERVAL, default 24h, first sweep immediately on boot) looking for disagreement between three things that are supposed to stay in step: workflow trigger config, Temporal schedules, and the workflow-run projection in Postgres.

It is a safety net. Drift is expected to be rare. That is exactly why it needs alerts — a control that fires once a quarter is a control nobody watches a dashboard for.

Six of its seven drift classes both detect and converge. One does not:

classdetectsconverges
missing_schedulea schedule-enabled config row with no Temporal schedulereissues it
paused_mismatchschedule paused state != configre-reconciles it
orphaned_schedulea wf-sched:* schedule with no owning config rowprunes it
stale_projectiona Temporal-owned run whose worker diedmarks it failed
orphaned_taska coordinator task whose backing run was deletedfails it closed
stranded_taska coordinator task whose backing run is terminalsettles it
orphaned_approvala pending governance approval for a settled/gone runnothing, by design

orphaned_approval is detect-only deliberately: voiding a clearance-gated governance row from an unattended sweeper is a mutation on the approvals ledger with no human in the loop. That decision belongs to an ADR and a founder, not to a sweeper. See the header of internal/workflow/wfreconcile/approvalreconcile.go.


Before you debug ANY of these: is the sweeper running?​

If WorkflowReconcileSweepsStalled is firing, treat every other alert in this file as meaningless rather than reassuring. A reconciler that has stopped emits no drift metrics at all, which looks identical to a healthy system.

kubectl -n platform logs deploy/reconciler --tail=200 | grep -i 'workflow drift reconciler'

Three things stop it, in decreasing order of likelihood:

  1. Another replica holds the advisory lock. The daemon exits rather than waits when upsquad.orgunit_reconciler is held elsewhere. Look for reconciler: acquiring advisory lock followed by a fatal.

  2. WF_RECONCILE_DISABLED=true. The kill-switch for this whole surface. It logs a WARN on startup when set.

  3. The OTel MeterProvider is gone. wfreconcile instruments through otel.Meter(...), which resolves against the global provider. If cmd/reconciler/main.go stops calling runtimemetrics.NewProvider, every wf_reconcile_* series vanishes while the daemon runs perfectly — and this alert then reports a stall that is not happening. That is the state the binary shipped in until #2495 — and it was unobservable at link level, not merely unexported: the base binary did not link internal/runtime/metrics, otel/sdk/metric or the Prometheus exporter at all.

    Two guards, pinning different things. Cite the right one.

    filepinscatches a deleted call site?
    meterprovider_callsite_test.gothe call site — a static AST oracle: run() must call <import>.NewProvideryes
    metrics_export_test.gothe mechanism — absent under the no-op provider, present after installing oneno (it builds its own provider)

    The second one reads like a guard for this and is not; that was caught in review of #2589 and is why the first exists.

Confirm what is actually exported before chasing anything else:

kubectl -n platform port-forward deploy/reconciler 9119:9119 &
curl -s localhost:9119/metrics | grep '^wf_reconcile_'

Expected shape (measured, not assumed — the OTel Prometheus exporter rewrites instrument names):

wf_reconcile_drift_detected_total{class="orphaned_approval",otel_scope_name="context-engine/workflow",...} 3

If wf_reconcile_* is absent but dualwrite_* is present, it is the MeterProvider, not the loop: dualwrite uses promauto and is unaffected by OTel wiring. That asymmetry is the fastest diagnostic in this runbook.


WorkflowReconcileOrphanedApprovalDetected (warning, ticket)​

Meaning. One or more governance_approvals rows are pending while the workflow run that opened them is terminal or hard-deleted (#2381).

Why it matters. The row renders in the Approvals queue as a live, clearance-gated card for work that no longer exists. An approver can answer it in good faith and the decision lands in the ledger against a dead run — an audit-coherence breach against the "every action is auditable" tenet, not a cosmetic one.

Why it is not a page. The breached state is durable, not time-critical. Paging at 03:00 cannot make an already-orphaned approval less orphaned, and the remediation is a daytime action with a real principal. It must never be silent, which is what it was until #2495. Route to a ticket queue.

What to do.

  1. Read the reconciler's WARN lines. One per row, each carrying org_id, approval_id, action_type, run_id and run_status:

    kubectl -n platform logs deploy/reconciler --since=26h \
    | grep 'orphaned pending approval'

    run_status="<run row deleted>" is the worse of the two breaches — no subsequent event can ever resolve that row.

  2. Void each id through the operator VoidApproval API. Do not UPDATE the table directly: the API writes the audited hop with a real principal, and the ledger is the audit trail.

  3. Check the direction of travel. Compare against the previous sweeps:

    sum by (class) (increase(wf_reconcile_drift_detected_total{class="orphaned_approval"}[26h]))

    Draining (count falling sweep over sweep) = you are clearing the pre-fix cohort; keep going. Growing = the #2381 interpreter fix has regressed and new leaks are being created. That is a P1 and a code fix, not a cleanup.

When it clears. Not on its own. The class converges nothing, so the alert clears only once a subsequent sweep finds no orphans — i.e. after you have voided them — plus up to 26h for the increase() window to drain.


WorkflowReconcileDriftNotConverging (warning)​

Meaning. A converging class detected drift and did not heal all of it in the same sweep. Detected minus converged should be 0 for every class except orphaned_approval.

Why it matters. The reconciler is finding the problem and failing to fix it. That is strictly worse than not looking, because the dashboard says the reconciler ran.

Convergence failures also bump workflow_schedule_reconcile_failures_total (shared with the WF-28 post-commit hook), so read that counter and the error log together.

Per class:

  • missing_schedule / paused_mismatch — the Temporal schedule reissue is failing. Check Temporal reachability and that TEMPORAL_NAMESPACE matches.
  • orphaned_schedule — the prune is failing. Expect occasional single-sweep noise here: Temporal's schedule List is a visibility index updated asynchronously after a delete, so one sweep can legitimately re-report a schedule it just pruned. The rule's for: 1h is sized to absorb that; a sustained firing is real.
  • stale_projection — the run-status heal is failing.
  • orphaned_task / stranded_task — the coordinator task terminalizer is refusing the transition. The graph guard is fail-closed by design: read the audit hop before assuming the reconciler is at fault.

WorkflowReconcileDriftSustained (warning)​

Meaning. A class has detected drift both now and two days ago.

Why it is separate from the alert above. That one says "the reconciler cannot fix it". This one says "the reconciler fixes it and something upstream keeps re-creating it". Convergence hides the underlying defect; the only remaining evidence is that the sweeper keeps having work to do. A safety net catching something every single day is a bug somewhere else.

Find the producer, not the sweeper. Check the WF-28 post-commit schedule hook and the interpreter's run-status projection before touching anything in internal/workflow/wfreconcile.


Editing these rules​

Run the full suite locally before pushing — it is ~5s once promtool is on PATH:

bash scripts/check-alert-rules.sh

It renders every PrometheusRule CRD into deploy/alerts/.rendered/ (gitignored), lints them, checks the beta-dev mirrors are in sync, and runs the promtool fixtures in deploy/alerts/tests/.

Two things to know before changing an expression:

  • This CRD is deliberately NOT mirrored into dev/prometheus/rules/. cmd/reconciler is not a service in docker-compose.dev.yml — it is published to GHCR and runs on GKE — so arming these on the beta-dev Prometheus would produce permanent absent() noise, and noise is how a real alert gets ignored. Same call, same reason, as dualwrite.yaml.
  • The fixture is not decoration. deploy/alerts/tests/wf-reconcile-drift_test.yaml pins four things that are each one edit away from a silently-dead rule: the or … * 0 zero-fill, the class!="orphaned_approval" exclusion, the offset 2d in the sustained rule, and the absent() arm of the stall rule. Every one of them was verified to go red when removed. If you delete an assertion because it "looks redundant", re-run that verification first.