Retrospective: Eight Failure Classes from Workflow Runtime v2
Date: 2026-08-03
Scope: milestone #28 — Workflow Runtime v2 (PRD #2081), WR2-1 through WR2-5, plus the pre-existing defects on main that this programme exposed
Filed by: #2380 (fold-in 2 of ARCH CONCERN #2346)
Sibling retro: Shelfware Pattern in Waves 1–3
Why this document exists
None of what follows is advice. Every class below is a defect that shipped or nearly shipped during milestone #28, and most of them recurred — one of them nine times. The programme's most uncomfortable finding is not any single bug; it is that each recurrence was introduced by someone who had already been told about the previous one, because the rule had been recorded as prose and not as a structure that fails.
This retro is the evidence surface: what happened, in which issue, and what measurement settled it. It deliberately does not restate the rules — see Where the rules live.
The programme proved its own thesis before this document existed
While compiling this retro (#2380) we found that two issues cite lessons that were never written:
- #2398 refs
`tasks/lessons.md` 2026-08-01 - #2392 refs
`tasks/lessons.md` (2026-07-31)
Neither entry exists. On origin/main at the time of writing,
git show origin/main:tasks/lessons.md | grep -cE '^##[^#]' returns 4 — two from
#1733 and two from #1987 — and the file was last touched by commit c362f052
(#2192 / #2228), well before either date. Both citations were confirmed independently
during review of PR #2434.
Two agents believed they had recorded a lesson, cited it as an authority in a P1 and a P2 issue, and the lesson was never committed. That is this document's own thesis happening in real time: a rule recorded as an intention rather than as a structure that fails does not exist, and nothing anywhere told anyone it was missing. The citations are corrected on the issues themselves and now point here.
The classes at a glance
Two separate questions, kept in separate columns because conflating them is how this table first went wrong: "reached prod" is a historical fact and never changes; "status" is as of the date below and needs re-checking. Where the current state has not actually been measured, the cell says unmeasured rather than guessing.
| # | Class | Grounding | Reached shipped main? | Status (2026-08-03) | Caught by |
|---|---|---|---|---|---|
| 1 | INERT-LOOP | #2346, PR #2344 | No — caught at review | Closed. PR #2344 closed unbuilt; re-cut as E-4a…E-4e | Architect reading the spec, not the code |
| 2 | PURITY-FAKE | #2384, #2382, #2344 | n/a — test harness | Open: 3 named fakes in interpreter_test.go (#2384). Counted, not estimated | Review of a sibling PR |
| 3 | Identity-dimension omission | #2398, #2396, #2397 | Yes | Exemplars fixed: PR #2441 made the lap a required parameter of the interpreter's hop-key derivation, so every audit hop key now carries it, and core#2397 closed the last named instance — the tool IdempotencyKey, now run:step@lap / run:step@lap#leg behind versionIdempotencyKeyLap. It was the LIVE one: a create_issue_comment in a loop body found lap 1's marker on lap 2. Population still open: PR #2441's enumeration found a NEW instance on the execution path — Tracked: #2442. Tracked: #2398 | Dimension enumeration during PR #2396 |
| 4 | Silent-drop dedupe | #2398 (P1) | Yes — 53% of audit hops in any loop | Exemplar fixed: PR #2441 closed both halves — the key carries the lap, and the audit dedupe now skips only a byte-identical row, reporting a divergent collision instead of committing it away. The narration event id (same hazard, ON CONFLICT DO NOTHING) was fixed with it. The 53% is measured, not estimated. Population UNMEASURED — no sweep of the repo's other ON CONFLICT DO NOTHING / check-and-skip writers has been run, and a content check cannot see a content-IDENTICAL collision at all. Tracked: #2398 | Counting, after presence assertions passed |
| 5 | Layered-rejection shadowing | #2392, PRs #2388/#2393 | Yes — the cadence floor was structurally unreachable when found | Exemplar fixed: pinned below the schema by TestValidateWait_CadenceFloorIsPinnedBelowTheSchema (validator_wait_observation_test.go:463). Population UNMEASURED — #2392's sweep (~40 constraints + 12 enums) has not been run, so the number of remaining instances is unknown, not "some" | Mutation sweep of guard bodies and call sites |
| 6 | Routing-edge descent | #2270 → PR #2274 | Yes — 2 of 9 escapes reached prod | Fixed: converged on stepExit in PR #2274 (merged 2026-07-28); #2270 closed; two CI guards now hold it | Human review, all nine times; CI never |
| 7 | Skip-reports-as-green | #2218 | Yes — the WR2-3 security suite was path-gated | Fixed: PR #2431 moved it into security-unit-tests.yml's path-independent lane; #2218 closed 2026-08-03. The class is unmeasured elsewhere — no audit of other path-filtered suites has been run | Reasoning about the path filter |
| 8 | Version gates unpinnable | #2312 → #2335 → #2342 | Yes — all 13 interpreter gates were unpinned | Fixed: #2312/#2335/#2342 all closed; Tier-1 AST registry + frozen pre-gate fixtures on main. One residual recorded in §8 (arm terminators pinned, operands not) | A survived mutation, reported rather than dropped |
Three of the eight (3, 4, 6) are the same shape at different altitudes: a set that must be enumerated exhaustively was instead enumerated by hand, per site, by whoever happened to be editing.
1. INERT-LOOP — a caller, but no reachable behaviour
Shelfware is code with no caller. This is code with a caller and no reachable behaviour, and the SHELFWARE CHECK passes cleanly on it.
runWaitPoll (WR2-4 poll mode, PR #2344) loops legs on a cadence, each leg
re-evaluating a CEL condition. It was called, wired, version-gated, fixtured and
replayed. It was also provably dead, and the proof is entirely in the
specification (ARCH CONCERN #2346):
- each leg's CEL input is
outputs[prev](wait.go→branchInput), orboot.Contextfor an entry wait; - while the wait is parked, nothing writes
outputs— the walk is one coroutine blocked inAwaitWithTimeout; the signal coroutines (drain.go) writest.events/st.cancelled/st.paused/st.answers; parallel branches write distinct keys andwalkBranchseedsprevper branch;boot.Contextis fixed at boot; - the evaluator is pure by mandate —
guardcel.NewEnv()declares exactly one variable (args) plus the CEL stdlib, with no custom functions, nonow(), no lookup (ADR-0021 decision 6, for Envoy portability).activities/branch.gosays so in its own words: "a pure function of its input (compile + eval, no DB, no clock, no rand)."
A pure function of a constant input is a constant. Every leg returns leg 1's boolean;
legs 2..N are dead work. For any validator-accepted definition the arm resolves on
leg 1 or runs to on_timeout.
The defect was in the LLD, not the PR. LLD #2253 §3.3 specified
exec(ActEvalBranch, {Expr: Condition, Input: pollSurface}) and never defined what
makes pollSurface differ between legs. The implementer built exactly that. The
architect's own words: "The PR's engineering is the best in the milestone; the thing
it was told to build is not buildable."
What the checklist was missing: can this loop's exit condition ever change?
Resolution: founder approved Option E (poll a governed tool_action observation);
PR #2344 was closed and the slice re-cut as E-4a…E-4e.
Correction to the popular framing
It is tempting to say "inert calls cluster — find one, sweep for siblings."
The issues do not support that for this class: runWaitPoll is the only INERT-LOOP
instance in the programme (n = 1). Clustering is very strongly evidenced for classes
2, 3, 5 and 6 — see the cross-cutting observation
— so the sweep instinct is right, but it is a property of the other classes, and the
INERT-LOOP defence is a per-loop question, not a sweep.
2. PURITY-FAKE — an input-independent fake cannot falsify a claim about its input
On PR #2344, eleven mutations went RED and the full replay corpus went green, over the provably dead loop above. That is the part worth understanding.
Both fakes keyed their scripted response off the call ordinal, ignoring their input
(interpreter_test.go; wait_fixture_gen_test.go). So both asserted "the loop
ITERATED" — a behaviour the documented-pure production evaluator cannot produce.
The claim actually under test was "leg N can differ from leg 1", which is purely a
claim about the input. The fake answered it by fiat. The fake did what production
could not.
The class is not confined to that PR. #2384 records three siblings in
internal/temporalwf/interpreter/interpreter_test.go, still open, after PR #2382
fixed a fourth (OpenQuestionGate minted one fixed ApprovalID regardless of input —
which would have made every test in the composition slice inert while reporting green):
| Fake | Defect |
|---|---|
OpenApprovalGate(_ context.Context, _ api.GateInput) | input discarded entirely; returns a fixed rec.gateApprovalID |
OpenUserPrompt | records the input, returns a fixed rec.promptApprovalID |
VoidWithDefault | records payload + reason but drops in.ApprovalID — the singular captures are last-write-wins, so a composed user_input timeout-default test cannot prove each lap resolved its own row |
Correct exemplars sit in the same file: QueryApprovalGate and QueryQuestionGate
stamp cp.ApprovalID = in.ApprovalID; resolveDefaultIDs records the id each
system-resolve targeted in call order; questionRowPerOpen mints a distinct row
per open, modelling production row identity.
Derive from the input, never from a global call counter. A counter makes the id depend on activity completion order, so two concurrently-parked branches swap ids run to run — measured on #2382.
The property that makes the rule worth having: force the fake to be pure and a test that needs leg 1 ≠ leg 2 must vary the input, which is exactly the property production has to provide. The test becomes unwritable exactly when the feature is unbuildable. That is the whole point, and it is why the rule matters most when the fake sits underneath a safety assertion — there, a fake that answers by fiat is manufacturing the safety evidence.
Register the real collaborator when it is pure
The sharpest form of the rule, and the one that would have killed #2344 outright:
guardcel.NewEnv() is pure by mandate and dependency-free (ADR-0021 decision 6 —
no DB, no clock, no rand, no custom functions). It was therefore free to register
in the test, and the fake bought nothing except the wrong answer.
Ordinal-scripted fakes pin arity and ordering; they cannot pin termination — which is precisely why eleven RED mutations and a green replay corpus did not stop a provably dead loop from reaching review. A single test registering the real evaluator would have shown every leg returning leg 1's boolean.
This only generalises where the collaborator really is pure. Where it is I/O-bound — a DB, an HTTP endpoint, the Temporal service — registering the real one in a workflow unit test is not reasonable, and the burden falls back on the input-derivation rule above and on naming the varying input.
3. Identity-dimension omission — every key must carry every dimension
A construct that repeats along (step, lap, leg, attempt) must carry every
dimension it repeats along in every identity key it participates in — idempotency
marker, audit hop, projection row, dedupe key.
Per #2398, the omission surfaced at six layers:
- the CEL surface (leg)
- the idempotency marker (leg)
- the audit hop (leg)
- the audit hop (lap)
- the
wait_observe_exhaustedhop (lap — found in flight) - two pre-existing repo-wide cases:
hopKeyandverifyHopKey
— plus #2397 at the platform level (IdempotencyKey was lap-blind, so a marker-checking
tool invoker returned lap 1's result on lap 2).
Both of those pre-existing cases were lap-blind on main until PR #2441, which made
the lap a required parameter of the derivation — so the compiler, not a reviewer,
now forbids minting a lap-blind hop key. The shipped shape is
hop:{step}:{event}:{lap}[:{ordinal}]
with the lap always emitted and the construct-local ordinal (leg | attempt | ord) last.
The platform-level case (#2397) has since been closed the same way and to the same
shape — run:step@lap for a tool_action and run:step@lap#leg for a poll
observation, composed from ONE spelling of the lap segment, always emitted, behind the
versionIdempotencyKeyLap GetVersion gate. Worth recording is that the fix needed a
version patch the audit-hop fixes did not: a hop key lives only in our own database,
but the idempotency key is embedded verbatim in the GITHUB ARTEFACT a create call
marks, so an in-flight run that changed shape mid-flight would stop finding its own
marker. Tracked: #2398, #2442.
The damning line, from #2398: "Each fix was made by someone who had just been told about the previous one."
Two technique notes that came out of it:
- Emit the segment unconditionally, so a positional value can never be ambiguous
between dimensions. #2396's
hop:{step}:wait_observe:{lap}:{leg}is the pattern — it emits both segments even when one is trivially zero, for exactly this reason. - Enumerate the keys and prove each dimension against the construct. The way this kept passing review was auditing the key list against itself — every key in the list looked consistent with every other, because they all shared the same omission.
4. Silent-drop dedupe — a missing dimension becomes invisible data loss
This is the highest-value entry in the retro. It was live on main, independent of
WR2 (#2398, escalated P2 → P1), and was fixed by PR #2441 — but the class is written
up in the tense it was found in, because the mechanism is what generalises.
Audit hop writes were check-and-skip, inside the chain lock
(internal/coordinator/audit_pg.go):
if hop.HopKey != "" {
var existingSeq int64
e := tx.QueryRow(ctx, `SELECT seq FROM coordinator_audit_log
WHERE task_id = $1 AND hop_key = $2 ...`, hop.TaskID, hop.HopKey).Scan(&existingSeq)
if e == nil {
return tx.Commit(ctx) // already appended under this key — idempotent no-op
}
...
}
hopKey omitted the lap (class 3),
so a hop emitted more than once across loop iterations produced one key and every later
row was discarded — no error, no log, just missing audit.
The measurement (architect, PR #2396 review 4834293531): a 3-lap loop emitted 19 hops
and silently lost 10 — 53%. The losses include govern verdicts on two nodes and the
agent result hop. That was core tenet 2 — "Every action is auditable" — broken for
any workflow containing a loop, from the day loops shipped until PR #2441.
Why it survived every test
Because the tests asserted presence. A presence assertion is satisfied by lap 1's
row and cannot see laps 2..N vanish. The falsifying measurement is a count: #2396's
4 observations across 2 laps collapsed to 2 distinct keys and the assertion read
should have 4 item(s), but has 2.
Test dedupe by counting emissions. Never by asserting presence.
Two findings that lower the cost of the fix
HopKeyis not an input toComputeAuditRowHash(verified: the canonical struct it marshals covers seq/step/role/action/verdict/reason/hashes/detail, and not the hop key). Adding the lap is therefore a pure key-derivation change — it perturbs no chain hash, so it is not an audit-format migration. No re-hashing, no historical rewrite.- Convention set on #2398 and applied in PR #2441: the shared
auditHopfamily (wait_open,govern,wait_resolved,wait_timeout*) was re-keyed once, centrally. Re-keyingwait_openin the poll arm alone would have given one event name two spellings depending on wait mode.
The design question, and how PR #2441 answered it
Is check-and-skip the right semantic for audit at all? A dropped audit row is not
obviously better than a duplicate one. ON CONFLICT DO NOTHING and check-and-skip
are the same hazard: they convert a key defect into silence.
The answer shipped is conditional skip: a key hit is a no-op only when the incoming hop re-derives the stored row's hash, and a divergent hit is reported rather than committed away. That is a real narrowing, and it is deliberately not sold as more than it is — a content check cannot see a collision between two content-IDENTICAL events, which is exactly why the key must carry every dimension. Neither half is sufficient alone. Tracked: #2398.
5. Layered-rejection shadowing — the inner layer is never reached
The predicate (use this one). A rule is shadowed iff the outer layer rejects a superset of what the rule rejects — rejection-set containment. It is not "the rule duplicates an outer constraint"; that framing flags nearly everything and is useless. Strictly-stronger rules stay reachable and pinned — the 200-leg ceiling, computed cross-field relations JSON Schema cannot express, and so on.
Detection is uniform across all three forms: delete the inner rule, run the suite. Green ⇒ shadowed.
Form A — schema → Go (#2392, from PR #2388's 34-mutation sweep)
The semantic cadence floor interval < minWaitPollIntervalSeconds (#2255) could be
deleted with the suite staying green. Its only test entered through
ValidateDefinition, where the JSON schema's "minimum": 30 rejects first. Worse than
untested: validateSemantics is private, has exactly one caller, is schema-first
and returns on failure — the rule was structurally unreachable in production.
Since fixed: the cadence floor is now pinned below the schema by
TestValidateWait_CadenceFloorIsPinnedBelowTheSchema
(internal/workflow/validator_wait_observation_test.go:463), which calls validateWait
directly rather than entering through ValidateDefinition.
The population is unmeasured, and that is the honest statement. #2392's sweep — ~40
minLength/minimum/maximum constraints plus 12 enums — has not been run. How many
of them shadow a Go rule is unknown. Not "some", not "several": unknown until someone
runs the deletion test against each.
Form B — Go gate → SQL predicate (PR #2393, E-4b)
Three SQL org_id predicates were unreachable because a Go gate (requireVisibleRun)
rejects first. Same containment shape, different medium. Retesting below the gate
pinned them — and the artifact case needed two orgs sharing a run_id to construct
an input the gate admits but the SQL must reject. The test had to be built to reach
past the shadowing layer.
A related-but-distinct find from the same PR: a gate satisfied by the gate below it — two sequential gates where the second's rejection makes the first unfalsifiable.
Form C — mutual shadowing (PR #2393 review 4833939651)
Same layer, same statement, two redundant conjuncts covering each other: an explicit
org_id = $2 predicate and the current_setting('app.org_id') RLS guard. Mutating
either one alone leaves the suite green. Deleting both is caught.
The fix is different in kind: you cannot pin these by testing below a layer, because there is no layer below. It needs a case where the two conjuncts disagree — here, param-org ≠ tx-GUC-org, which one rejects and the other admits. A single such test kills both mutants.
For any redundant pair of conditions guarding the same thing, a test must exist where the two disagree. Otherwise defence-in-depth is unverifiable by construction — and silently becomes defence-in-one the first time someone simplifies.
The direction that matters most
The dangerous move is the outer layer being relaxed, not the inner rule deleted. A
test asserting only "validation failed" cannot say which layer rejected, so it pins
neither. Loosen the schema and the Go rule silently becomes the sole untested gate;
delete the Go rule and the schema silently becomes the sole untested gate.
Audit the pair, not either half — and assert which layer rejected.
Two technique notes worth carrying: assert on error identity, not a substring probe
(#2388's hardened assertion killed an inlined look-alike that satisfied a substring —
and the architect's own proposed seam would have passed that mutant); and make a seam
t.Fatal if the case under test is already admitted, so a later slice cannot
silently render the test vacuous.
6. Routing-edge descent — N hand-maintained copies of one set
Nine escape instances across four PRs. Every one caught by human review, none by CI.
Two reached shipped main (#2270, escalated to P1 — both since closed by PR #2274,
merged 2026-07-28; stepExit is on main today. The two below are stated as they
stood when #2270 was filed):
- P4 —
user_input.on_timeout_nextbypasses the parallel join: a step inside a parallel branch escapes to a post-join successor without joining. - P5 — the same edge bypasses the loop exit: escapes the loop region without going through the loop's exit path.
These are containment/correctness bugs, not hop under-counts: a governed run could
route around a join or a loop boundary. Probes P3b/P3e reproduced the residual region
escape with plain agent_action / approval_gate — no WR2 node type needed. The
gap was generic and pre-existing.
The actual defect
Five independent whole-graph walks each hand-maintained their own edge enumeration:
cycleEdgeKeys · hopCalc · validateRegionScopes · regionSuccessors · loopBodyTails.
Nothing forced them to agree, so the #2193 "descend every routing edge" fence was a convention, not a structure: a new node type got descended by whichever passes its author happened to think of. PR #2271 descended 3 of 5 and looked complete.
The clearest evidence that a convention cannot hold this: loopBodyTails was found
blind twice, independently, in the same week, by two different agents, on two
different node types (#2269 wait, #2271 question) — and in both cases only while
chasing a different reported defect.
What actually closed it (PR #2274)
One vocabulary — type stepExit struct{ Key, Target string; CarriesEnvelope bool } —
with three accessors (stepExits/stepExitTargets, stepContainmentTargets,
stepRegionTargets), every walk consuming those and nothing else. Routing, containment
and envelope-bearing stay three separate sets by explicit architect instruction.
Then the part that makes instance #10 impossible — two CI guards:
TestCIGuard_NoHandRolledCycleEdgeKeysWalks— AST-parses the package and fails when any function outside a named, documented allow-list ranges overcycleEdgeKeys. The allow-list is also checked for staleness in reverse.TestCIGuard_EverySuccessorBearingStepFieldIsReachable— reflects overdsl.Step, detects every field whose JSON name looks like a successor (next,*_next,on_*,body,branches— deliberately over-broad), and fails unless it is either declared an edge and actually resolvable throughstepRegionTargets, or recorded as a policy word with a reason.
Mutation M1 — dropping the single line case "user_input": add("on_timeout_next") —
reddened six tests. That is the proof the centralisation holds: one declaration,
every consumer inherits.
The trap inside the fix
on_timeout on a user_input is a policy word (fail|default|route), not an edge.
On a wait, the identically-named field is a branch name and is an edge.
Admitting it universally would make every walk chase a step named "route". Hence the
guard's second arm: a successor-looking field may be excluded, but only by being
recorded as a policy word with a reason — never by silence.
7. Skip-reports-as-green — a required check satisfied by not running
A path-filtered CI job that does not fire reports success, not "skipped", at the
roll-up level. A required status check is therefore satisfied by a job that never ran.
The milestone-#28 instance (#2218, since fixed — see below): the WR2-3 sandbox/workspace
security-assertion suite runs through db.yml's Migrate-Up/Migrate-Down integration
lane, gated on the internal/workflow/** + cmd/agent-orchestrator/** path filters. A
dependency-only change — e.g. to internal/mcp/egress — that could weaken the
sandbox's egress/isolation guarantees will not re-fire the suite, because it does
not touch the filtered paths. The PR went green with the security assertions unexecuted.
Fixed: PR #2431 moved the two packages into security-unit-tests.yml, which triggers
on every PR with no top-level paths: filter, so the suite now runs regardless of which
subtree moved. #2218 closed 2026-08-03.
The class is not measured elsewhere. No audit of the repo's other path-filtered jobs
has been run, so how many other invariant-bearing suites sit behind a changes filter is
unknown — which is the reason the check below is stated as a question to ask, not as a
count of known instances.
Two consequences, and the second is the one that costs hours:
- Gating: an invariant/security suite must gate path-independently of the change surface. Path filters are a cost optimisation for ordinary lanes; on a suite whose job is to prove an invariant, they are a hole.
- Diagnosis: comparing a skipped run against a real run is not a comparison. Read
the per-job conclusion (
skippedvssuccess), never the roll-up, before concluding "the same suite passed on the other branch".
8. Version gates unpinnable — the oracle excludes the thing you want to pin
No replay fixture we have, or could record, can falsify the deletion of a
workflow.GetVersion call. Deleting a gate leaves the entire 32-fixture replay corpus
green. All 13 gates in the interpreter rested on reasoning alone (#2312).
This is deliberate SDK behaviour, not a corpus gap. In go.temporal.io/sdk@v1.46.0:
skipDeterministicCheckForCommand excludes a RECORD_MARKER command named
versionMarkerName; skipDeterministicCheckForEvent excludes the matching
MARKER_RECORDED event; and incrementNextCommandEventIDIfVersionMarker skips the
marker and its paired UpsertWorkflowSearchAttributes event when matching command
positions. The exclusion is what lets you delete a GetVersion call once old runs
drain. The side effect is that the marker is invisible to our only determinism oracle.
Surfaced because a survived mutation was reported rather than quietly dropped (on #2258 / PR #2310). That reporting discipline is the reason this class is in the retro at all.
What actually pins a gate (#2335)
- Tier 1 — AST gate registry.
internal/temporalwf/version_gates_test.goenumerates everyworkflow.GetVersioncall with its change-id, const,maxSupported, and its exact lexical scope path, e.g.func dispatchStep > switch(step.Type) > case(dsl.StepWait) > if(step.WaitMode == waitModeSignal && GETVERSION >= 1). The scope path is what catches mis-scoping rather than mere deletion: dropping a guard predicate (widening) or adding one (narrowing) changes the fingerprint, while a cosmetic reformat does not.versionContractCheckhas 4 sites and the comparison is a multiset, so deleting 1 of 4 is red. - Tier 2 — frozen pre-gate fixtures, recorded against a binary whose DefaultVersion
arm is live, for the gates where a live in-flight run could hold the pre-gate shape.
The generator asserts the pre-gate property (no marker, no post-gate command) so
running it against patched code fails loudly instead of quietly writing a useless
history. Hand-editing an existing fixture to strip its marker is not a shortcut — the
replayer fails with
missing history events, expectedNextEventID=…, because event ids must stay contiguous and are cross-referenced throughout. - Tier 3 — the rule that protects Tier 2. Bulk fixture regeneration is a silent
gate-unpinning operation. Every
*_pregate_*fixture is excluded from every regenerate path and frozen at_v0by definition.
Three findings nobody expected
- Two gates were already pinned by accident and nobody knew.
wf-run-log-narrationgoes red on 11 pre-#1987 histories;wf16-cancel-voids-pending-approvaloncancel_v1.json. Nothing recorded it — and a routine regeneration oflinear_dag_v1.jsonorcancel_v1.jsonwould have silently destroyed both. - A third category exists: reachable but not falsifiable. The replayer compares the
command stream, not payloads or timer durations.
lbe-b3-user-input-asked-byonly changes an activity's input payload, so it replays green when deleted; the const's own stated rationale was wrong about why. Recorded as a category rather than pretended away. Likewisegate_approve_v1/gate_reject_v1do not pinwf-gate-repoll— a 15s poll leg matches a 24h wait command-for-command, and the signal resolves before any re-poll command is emitted. The timer had to elapse. - The fourth evasion (#2339 → #2342): one token.
if GetVersion(...) < 1 { old; return }was voided by deleting thereturn— areturnis not a call, so the callee-set key stayed byte-identical and Tier 1 passed. The fix renders each arm as callees + terminator (return/break/continue/goto/fallthrough/panic, label included —break:legs≠break); 11 of 17 registry entries gained+return. Terminators that are calls (os.Exit, an always-panicking helper) are deliberately not enumerated — they already appear in the callee set, so swapping one in changes the key twice over. - The author then defeated his own fix and closed it in the same PR. The same gate
written as a
switchcase rendersgated=[] other=[], because arm extraction only reads if-conditions — so a gate authored in that shape is born with no arm content and every edit inside it is invisible. Converting an existing gate is caught (the ancestor path changes); a new one would not be. Tier 1 now rejects the shape outright with an actionable message. 52 mutations, 0 unexpected. - Residual, recorded not closed: an arm's terminator is pinned but not its
operands —
return ok→return falsekeeps the key green (measured). The gate is not voided (the arm still does not fall through) but its result changes.
The generalisation
A green suite is not evidence that a construct exists, if the oracle deliberately excludes that construct. When it does, no amount of that oracle will pin it — you need a second oracle at a different level (here: structural/AST over source, beside behavioural over history).
The observation that generalises
Six of the eight classes share one root: a set that had to be enumerated exhaustively was enumerated by hand, per site, by whoever was editing.
| Class | The set | Copies |
|---|---|---|
| 3 / 4 | dimensions a key must carry | one per key-minting site |
| 5 | layers that can reject | implicit, never written down |
| 6 | routing edges | five walks |
| 8 | version gates | thirteen, pinned by reasoning |
In every case the failure was invisible to review because each site was internally consistent — a hand-maintained list audited against itself always agrees with itself. The only things that ended a recurrence were convergence to one declaration plus a CI guard that fails when a member is added without registration (class 6), or a second oracle at a different level (class 8). Prose rules did not: classes 3 and 6 each recurred after the rule had been stated and acknowledged.
The clustering corollary, which is well evidenced for classes 2, 3, 5 and 6 (and not for class 1 — see the correction there): when you find one instance, the sibling instances already exist. #2382 fixed one input-independent fake and three siblings sat in the same file. #2270 fixed one routing edge and eight more instances followed. Budget the sweep at discovery time, not after the second recurrence.
Where the rules live
Deliberately split so no rule has two homes:
| Surface | Holds | Path |
|---|---|---|
| This retro | the evidence — what happened, in which issue, what measurement settled it | docs/retrospectives/2026-08-03-wr2-failure-classes.md |
| Lessons | the author-side construction rules — what to do while writing the code or the test | tasks/lessons.md |
| Reviewer checklist | the review-time questions — what to ask of a diff, and what evidence refutes the concern | .claude/agents/principal-architect.md, beside the SHELFWARE CHECK |
Each surface links to the others rather than restating them. If a rule needs changing, change it in its one home.
Refs
#2380 (this write-up) · #2346 · #2384 · #2398 · #2396 · #2397 · #2392 · #2388 · #2393 · #2270 · #2274 · #2271 · #2269 · #2218 · #2312 · #2335 · #2339 · #2342 · PR #2344 (closed) · LLD #2253 · HLD #2223 · PRD #2081 §4.4 · Tracker #2103