Runbook — WF-31 workflow drift reconciler alerts
Target of the runbook_url on every alert in
deploy/alerts/wf-reconcile-drift.yaml.
Owner: platform / devops. Issue: #2495.
What the reconciler is, and what it is not
internal/workflow/wfreconcile, run by the cmd/reconciler daemon, sweeps every
tenant once a day (WF_RECONCILE_INTERVAL, default 24h, first sweep immediately
on boot) looking for disagreement between three things that are supposed to stay
in step: workflow trigger config, Temporal schedules, and the workflow-run
projection in Postgres.
It is a safety net. Drift is expected to be rare. That is exactly why it needs alerts — a control that fires once a quarter is a control nobody watches a dashboard for.
Six of its seven drift classes both detect and converge. One does not:
| class | detects | converges |
|---|---|---|
missing_schedule | a schedule-enabled config row with no Temporal schedule | reissues it |
paused_mismatch | schedule paused state != config | re-reconciles it |
orphaned_schedule | a wf-sched:* schedule with no owning config row | prunes it |
stale_projection | a Temporal-owned run whose worker died | marks it failed |
orphaned_task | a coordinator task whose backing run was deleted | fails it closed |
stranded_task | a coordinator task whose backing run is terminal | settles it |
orphaned_approval | a pending governance approval for a settled/gone run | nothing, by design |
orphaned_approval is detect-only deliberately: voiding a clearance-gated
governance row from an unattended sweeper is a mutation on the approvals ledger
with no human in the loop. That decision belongs to an ADR and a founder, not to
a sweeper. See the header of internal/workflow/wfreconcile/approvalreconcile.go.
Before you debug ANY of these: is the sweeper running?
If WorkflowReconcileSweepsStalled is firing, treat every other alert in
this file as meaningless rather than reassuring. A reconciler that has
stopped emits no drift metrics at all, which looks identical to a healthy
system.
kubectl -n platform logs deploy/reconciler --tail=200 | grep -i 'workflow drift reconciler'
Three things stop it, in decreasing order of likelihood:
-
Another replica holds the advisory lock. The daemon exits rather than waits when
upsquad.orgunit_reconcileris held elsewhere. Look forreconciler: acquiring advisory lockfollowed by a fatal. -
WF_RECONCILE_DISABLED=true. The kill-switch for this whole surface. It logs a WARN on startup when set. -
The OTel MeterProvider is gone.
wfreconcileinstruments throughotel.Meter(...), which resolves against the global provider. Ifcmd/reconciler/main.gostops callingruntimemetrics.NewProvider, everywf_reconcile_*series vanishes while the daemon runs perfectly — and this alert then reports a stall that is not happening. That is the state the binary shipped in until #2495 — and it was unobservable at link level, not merely unexported: the base binary did not linkinternal/runtime/metrics,otel/sdk/metricor the Prometheus exporter at all.Two guards, pinning different things. Cite the right one.
file pins catches a deleted call site? meterprovider_callsite_test.gothe call site — a static AST oracle: run()must call<import>.NewProvideryes metrics_export_test.gothe mechanism — absent under the no-op provider, present after installing one no (it builds its own provider) The second one reads like a guard for this and is not; that was caught in review of #2589 and is why the first exists.
Confirm what is actually exported before chasing anything else:
kubectl -n platform port-forward deploy/reconciler 9119:9119 &
curl -s localhost:9119/metrics | grep '^wf_reconcile_'
Expected shape (measured, not assumed — the OTel Prometheus exporter rewrites instrument names):
wf_reconcile_drift_detected_total{class="orphaned_approval",otel_scope_name="context-engine/workflow",...} 3
If wf_reconcile_* is absent but dualwrite_* is present, it is the
MeterProvider, not the loop: dualwrite uses promauto and is unaffected by
OTel wiring. That asymmetry is the fastest diagnostic in this runbook.
WorkflowReconcileOrphanedApprovalDetected (warning, ticket)
Meaning. One or more governance_approvals rows are pending while the
workflow run that opened them is terminal or hard-deleted (#2381).
Why it matters. The row renders in the Approvals queue as a live, clearance-gated card for work that no longer exists. An approver can answer it in good faith and the decision lands in the ledger against a dead run — an audit-coherence breach against the "every action is auditable" tenet, not a cosmetic one.
Why it is not a page. The breached state is durable, not time-critical. Paging at 03:00 cannot make an already-orphaned approval less orphaned, and the remediation is a daytime action with a real principal. It must never be silent, which is what it was until #2495. Route to a ticket queue.
What to do.
-
Read the reconciler's WARN lines. One per row, each carrying
org_id,approval_id,action_type,run_idandrun_status:kubectl -n platform logs deploy/reconciler --since=26h \| grep 'orphaned pending approval'run_status="<run row deleted>"is the worse of the two breaches — no subsequent event can ever resolve that row. -
Void each id through the operator
VoidApprovalAPI. Do not UPDATE the table directly: the API writes the audited hop with a real principal, and the ledger is the audit trail. -
Check the direction of travel. Compare against the previous sweeps:
sum by (class) (increase(wf_reconcile_drift_detected_total{class="orphaned_approval"}[26h]))Draining (count falling sweep over sweep) = you are clearing the pre-fix cohort; keep going. Growing = the #2381 interpreter fix has regressed and new leaks are being created. That is a P1 and a code fix, not a cleanup.
When it clears. Not on its own. The class converges nothing, so the alert
clears only once a subsequent sweep finds no orphans — i.e. after you have voided
them — plus up to 26h for the increase() window to drain.
WorkflowReconcileDriftNotConverging (warning)
Meaning. A converging class detected drift and did not heal all of it in the
same sweep. Detected minus converged should be 0 for every class except
orphaned_approval.
Why it matters. The reconciler is finding the problem and failing to fix it. That is strictly worse than not looking, because the dashboard says the reconciler ran.
Convergence failures also bump workflow_schedule_reconcile_failures_total
(shared with the WF-28 post-commit hook), so read that counter and the error log
together.
Per class:
missing_schedule/paused_mismatch— the Temporal schedule reissue is failing. Check Temporal reachability and thatTEMPORAL_NAMESPACEmatches.orphaned_schedule— the prune is failing. Expect occasional single-sweep noise here: Temporal's scheduleListis a visibility index updated asynchronously after a delete, so one sweep can legitimately re-report a schedule it just pruned. The rule'sfor: 1his sized to absorb that; a sustained firing is real.stale_projection— the run-status heal is failing.orphaned_task/stranded_task— the coordinator task terminalizer is refusing the transition. The graph guard is fail-closed by design: read the audit hop before assuming the reconciler is at fault.
WorkflowReconcileDriftSustained (warning)
Meaning. A class has detected drift both now and two days ago.
Why it is separate from the alert above. That one says "the reconciler cannot fix it". This one says "the reconciler fixes it and something upstream keeps re-creating it". Convergence hides the underlying defect; the only remaining evidence is that the sweeper keeps having work to do. A safety net catching something every single day is a bug somewhere else.
Find the producer, not the sweeper. Check the WF-28 post-commit schedule
hook and the interpreter's run-status projection before touching anything in
internal/workflow/wfreconcile.
Editing these rules
Run the full suite locally before pushing — it is ~5s once promtool is on
PATH:
bash scripts/check-alert-rules.sh
It renders every PrometheusRule CRD into deploy/alerts/.rendered/
(gitignored), lints them, checks the beta-dev mirrors are in sync, and runs the
promtool fixtures in deploy/alerts/tests/.
Two things to know before changing an expression:
- This CRD is deliberately NOT mirrored into
dev/prometheus/rules/.cmd/reconcileris not a service indocker-compose.dev.yml— it is published to GHCR and runs on GKE — so arming these on the beta-dev Prometheus would produce permanentabsent()noise, and noise is how a real alert gets ignored. Same call, same reason, asdualwrite.yaml. - The fixture is not decoration.
deploy/alerts/tests/wf-reconcile-drift_test.yamlpins four things that are each one edit away from a silently-dead rule: theor … * 0zero-fill, theclass!="orphaned_approval"exclusion, theoffset 2din the sustained rule, and theabsent()arm of the stall rule. Every one of them was verified to go red when removed. If you delete an assertion because it "looks redundant", re-run that verification first.