Skip to main content

Retrospective: Eight Failure Classes from Workflow Runtime v2

Date: 2026-08-03 Scope: milestone #28 — Workflow Runtime v2 (PRD #2081), WR2-1 through WR2-5, plus the pre-existing defects on main that this programme exposed Filed by: #2380 (fold-in 2 of ARCH CONCERN #2346) Sibling retro: Shelfware Pattern in Waves 1–3


Why this document exists​

None of what follows is advice. Every class below is a defect that shipped or nearly shipped during milestone #28, and most of them recurred — one of them nine times. The programme's most uncomfortable finding is not any single bug; it is that each recurrence was introduced by someone who had already been told about the previous one, because the rule had been recorded as prose and not as a structure that fails.

This retro is the evidence surface: what happened, in which issue, and what measurement settled it. It deliberately does not restate the rules — see Where the rules live.

The programme proved its own thesis before this document existed​

While compiling this retro (#2380) we found that two issues cite lessons that were never written:

  • #2398 refs `tasks/lessons.md` 2026-08-01
  • #2392 refs `tasks/lessons.md` (2026-07-31)

Neither entry exists. On origin/main at the time of writing, git show origin/main:tasks/lessons.md | grep -cE '^##[^#]' returns 4 — two from #1733 and two from #1987 — and the file was last touched by commit c362f052 (#2192 / #2228), well before either date. Both citations were confirmed independently during review of PR #2434.

Two agents believed they had recorded a lesson, cited it as an authority in a P1 and a P2 issue, and the lesson was never committed. That is this document's own thesis happening in real time: a rule recorded as an intention rather than as a structure that fails does not exist, and nothing anywhere told anyone it was missing. The citations are corrected on the issues themselves and now point here.

The classes at a glance​

Two separate questions, kept in separate columns because conflating them is how this table first went wrong: "reached prod" is a historical fact and never changes; "status" is as of the date below and needs re-checking. Where the current state has not actually been measured, the cell says unmeasured rather than guessing.

#ClassGroundingReached shipped main?Status (2026-08-03)Caught by
1INERT-LOOP#2346, PR #2344No — caught at reviewClosed. PR #2344 closed unbuilt; re-cut as E-4a…E-4eArchitect reading the spec, not the code
2PURITY-FAKE#2384, #2382, #2344n/a — test harnessOpen: 3 named fakes in interpreter_test.go (#2384). Counted, not estimatedReview of a sibling PR
3Identity-dimension omission#2398, #2396, #2397YesExemplars fixed: PR #2441 made the lap a required parameter of the interpreter's hop-key derivation, so every audit hop key now carries it, and core#2397 closed the last named instance — the tool IdempotencyKey, now run:step@lap / run:step@lap#leg behind versionIdempotencyKeyLap. It was the LIVE one: a create_issue_comment in a loop body found lap 1's marker on lap 2. Population still open: PR #2441's enumeration found a NEW instance on the execution path — Tracked: #2442. Tracked: #2398Dimension enumeration during PR #2396
4Silent-drop dedupe#2398 (P1)Yes — 53% of audit hops in any loopExemplar fixed: PR #2441 closed both halves — the key carries the lap, and the audit dedupe now skips only a byte-identical row, reporting a divergent collision instead of committing it away. The narration event id (same hazard, ON CONFLICT DO NOTHING) was fixed with it. The 53% is measured, not estimated. Population UNMEASURED — no sweep of the repo's other ON CONFLICT DO NOTHING / check-and-skip writers has been run, and a content check cannot see a content-IDENTICAL collision at all. Tracked: #2398Counting, after presence assertions passed
5Layered-rejection shadowing#2392, PRs #2388/#2393Yes — the cadence floor was structurally unreachable when foundExemplar fixed: pinned below the schema by TestValidateWait_CadenceFloorIsPinnedBelowTheSchema (validator_wait_observation_test.go:463). Population UNMEASURED — #2392's sweep (~40 constraints + 12 enums) has not been run, so the number of remaining instances is unknown, not "some"Mutation sweep of guard bodies and call sites
6Routing-edge descent#2270 → PR #2274Yes — 2 of 9 escapes reached prodFixed: converged on stepExit in PR #2274 (merged 2026-07-28); #2270 closed; two CI guards now hold itHuman review, all nine times; CI never
7Skip-reports-as-green#2218Yes — the WR2-3 security suite was path-gatedFixed: PR #2431 moved it into security-unit-tests.yml's path-independent lane; #2218 closed 2026-08-03. The class is unmeasured elsewhere — no audit of other path-filtered suites has been runReasoning about the path filter
8Version gates unpinnable#2312 → #2335 → #2342Yes — all 13 interpreter gates were unpinnedFixed: #2312/#2335/#2342 all closed; Tier-1 AST registry + frozen pre-gate fixtures on main. One residual recorded in §8 (arm terminators pinned, operands not)A survived mutation, reported rather than dropped

Three of the eight (3, 4, 6) are the same shape at different altitudes: a set that must be enumerated exhaustively was instead enumerated by hand, per site, by whoever happened to be editing.


1. INERT-LOOP — a caller, but no reachable behaviour​

Shelfware is code with no caller. This is code with a caller and no reachable behaviour, and the SHELFWARE CHECK passes cleanly on it.

runWaitPoll (WR2-4 poll mode, PR #2344) loops legs on a cadence, each leg re-evaluating a CEL condition. It was called, wired, version-gated, fixtured and replayed. It was also provably dead, and the proof is entirely in the specification (ARCH CONCERN #2346):

  • each leg's CEL input is outputs[prev] (wait.go → branchInput), or boot.Context for an entry wait;
  • while the wait is parked, nothing writes outputs — the walk is one coroutine blocked in AwaitWithTimeout; the signal coroutines (drain.go) write st.events/st.cancelled/st.paused/st.answers; parallel branches write distinct keys and walkBranch seeds prev per branch; boot.Context is fixed at boot;
  • the evaluator is pure by mandate — guardcel.NewEnv() declares exactly one variable (args) plus the CEL stdlib, with no custom functions, no now(), no lookup (ADR-0021 decision 6, for Envoy portability). activities/branch.go says so in its own words: "a pure function of its input (compile + eval, no DB, no clock, no rand)."

A pure function of a constant input is a constant. Every leg returns leg 1's boolean; legs 2..N are dead work. For any validator-accepted definition the arm resolves on leg 1 or runs to on_timeout.

The defect was in the LLD, not the PR. LLD #2253 §3.3 specified exec(ActEvalBranch, {Expr: Condition, Input: pollSurface}) and never defined what makes pollSurface differ between legs. The implementer built exactly that. The architect's own words: "The PR's engineering is the best in the milestone; the thing it was told to build is not buildable."

What the checklist was missing: can this loop's exit condition ever change? Resolution: founder approved Option E (poll a governed tool_action observation); PR #2344 was closed and the slice re-cut as E-4a…E-4e.

It is tempting to say "inert calls cluster — find one, sweep for siblings." The issues do not support that for this class: runWaitPoll is the only INERT-LOOP instance in the programme (n = 1). Clustering is very strongly evidenced for classes 2, 3, 5 and 6 — see the cross-cutting observation — so the sweep instinct is right, but it is a property of the other classes, and the INERT-LOOP defence is a per-loop question, not a sweep.


2. PURITY-FAKE — an input-independent fake cannot falsify a claim about its input​

On PR #2344, eleven mutations went RED and the full replay corpus went green, over the provably dead loop above. That is the part worth understanding.

Both fakes keyed their scripted response off the call ordinal, ignoring their input (interpreter_test.go; wait_fixture_gen_test.go). So both asserted "the loop ITERATED" — a behaviour the documented-pure production evaluator cannot produce. The claim actually under test was "leg N can differ from leg 1", which is purely a claim about the input. The fake answered it by fiat. The fake did what production could not.

The class is not confined to that PR. #2384 records three siblings in internal/temporalwf/interpreter/interpreter_test.go, still open, after PR #2382 fixed a fourth (OpenQuestionGate minted one fixed ApprovalID regardless of input — which would have made every test in the composition slice inert while reporting green):

FakeDefect
OpenApprovalGate(_ context.Context, _ api.GateInput)input discarded entirely; returns a fixed rec.gateApprovalID
OpenUserPromptrecords the input, returns a fixed rec.promptApprovalID
VoidWithDefaultrecords payload + reason but drops in.ApprovalID — the singular captures are last-write-wins, so a composed user_input timeout-default test cannot prove each lap resolved its own row

Correct exemplars sit in the same file: QueryApprovalGate and QueryQuestionGate stamp cp.ApprovalID = in.ApprovalID; resolveDefaultIDs records the id each system-resolve targeted in call order; questionRowPerOpen mints a distinct row per open, modelling production row identity.

Derive from the input, never from a global call counter. A counter makes the id depend on activity completion order, so two concurrently-parked branches swap ids run to run — measured on #2382.

The property that makes the rule worth having: force the fake to be pure and a test that needs leg 1 ≠ leg 2 must vary the input, which is exactly the property production has to provide. The test becomes unwritable exactly when the feature is unbuildable. That is the whole point, and it is why the rule matters most when the fake sits underneath a safety assertion — there, a fake that answers by fiat is manufacturing the safety evidence.

Register the real collaborator when it is pure​

The sharpest form of the rule, and the one that would have killed #2344 outright: guardcel.NewEnv() is pure by mandate and dependency-free (ADR-0021 decision 6 — no DB, no clock, no rand, no custom functions). It was therefore free to register in the test, and the fake bought nothing except the wrong answer.

Ordinal-scripted fakes pin arity and ordering; they cannot pin termination — which is precisely why eleven RED mutations and a green replay corpus did not stop a provably dead loop from reaching review. A single test registering the real evaluator would have shown every leg returning leg 1's boolean.

This only generalises where the collaborator really is pure. Where it is I/O-bound — a DB, an HTTP endpoint, the Temporal service — registering the real one in a workflow unit test is not reasonable, and the burden falls back on the input-derivation rule above and on naming the varying input.


3. Identity-dimension omission — every key must carry every dimension​

A construct that repeats along (step, lap, leg, attempt) must carry every dimension it repeats along in every identity key it participates in — idempotency marker, audit hop, projection row, dedupe key.

Per #2398, the omission surfaced at six layers:

  1. the CEL surface (leg)
  2. the idempotency marker (leg)
  3. the audit hop (leg)
  4. the audit hop (lap)
  5. the wait_observe_exhausted hop (lap — found in flight)
  6. two pre-existing repo-wide cases: hopKey and verifyHopKey

— plus #2397 at the platform level (IdempotencyKey was lap-blind, so a marker-checking tool invoker returned lap 1's result on lap 2).

Both of those pre-existing cases were lap-blind on main until PR #2441, which made the lap a required parameter of the derivation — so the compiler, not a reviewer, now forbids minting a lap-blind hop key. The shipped shape is

hop:{step}:{event}:{lap}[:{ordinal}]

with the lap always emitted and the construct-local ordinal (leg | attempt | ord) last.

The platform-level case (#2397) has since been closed the same way and to the same shape — run:step@lap for a tool_action and run:step@lap#leg for a poll observation, composed from ONE spelling of the lap segment, always emitted, behind the versionIdempotencyKeyLap GetVersion gate. Worth recording is that the fix needed a version patch the audit-hop fixes did not: a hop key lives only in our own database, but the idempotency key is embedded verbatim in the GITHUB ARTEFACT a create call marks, so an in-flight run that changed shape mid-flight would stop finding its own marker. Tracked: #2398, #2442.

The damning line, from #2398: "Each fix was made by someone who had just been told about the previous one."

Two technique notes that came out of it:

  • Emit the segment unconditionally, so a positional value can never be ambiguous between dimensions. #2396's hop:{step}:wait_observe:{lap}:{leg} is the pattern — it emits both segments even when one is trivially zero, for exactly this reason.
  • Enumerate the keys and prove each dimension against the construct. The way this kept passing review was auditing the key list against itself — every key in the list looked consistent with every other, because they all shared the same omission.

4. Silent-drop dedupe — a missing dimension becomes invisible data loss​

This is the highest-value entry in the retro. It was live on main, independent of WR2 (#2398, escalated P2 → P1), and was fixed by PR #2441 — but the class is written up in the tense it was found in, because the mechanism is what generalises.

Audit hop writes were check-and-skip, inside the chain lock (internal/coordinator/audit_pg.go):

if hop.HopKey != "" {
var existingSeq int64
e := tx.QueryRow(ctx, `SELECT seq FROM coordinator_audit_log
WHERE task_id = $1 AND hop_key = $2 ...`, hop.TaskID, hop.HopKey).Scan(&existingSeq)
if e == nil {
return tx.Commit(ctx) // already appended under this key — idempotent no-op
}
...
}

hopKey omitted the lap (class 3), so a hop emitted more than once across loop iterations produced one key and every later row was discarded — no error, no log, just missing audit.

The measurement (architect, PR #2396 review 4834293531): a 3-lap loop emitted 19 hops and silently lost 10 — 53%. The losses include govern verdicts on two nodes and the agent result hop. That was core tenet 2 — "Every action is auditable" — broken for any workflow containing a loop, from the day loops shipped until PR #2441.

Why it survived every test​

Because the tests asserted presence. A presence assertion is satisfied by lap 1's row and cannot see laps 2..N vanish. The falsifying measurement is a count: #2396's 4 observations across 2 laps collapsed to 2 distinct keys and the assertion read should have 4 item(s), but has 2.

Test dedupe by counting emissions. Never by asserting presence.

Two findings that lower the cost of the fix​

  • HopKey is not an input to ComputeAuditRowHash (verified: the canonical struct it marshals covers seq/step/role/action/verdict/reason/hashes/detail, and not the hop key). Adding the lap is therefore a pure key-derivation change — it perturbs no chain hash, so it is not an audit-format migration. No re-hashing, no historical rewrite.
  • Convention set on #2398 and applied in PR #2441: the shared auditHop family (wait_open, govern, wait_resolved, wait_timeout*) was re-keyed once, centrally. Re-keying wait_open in the poll arm alone would have given one event name two spellings depending on wait mode.

The design question, and how PR #2441 answered it​

Is check-and-skip the right semantic for audit at all? A dropped audit row is not obviously better than a duplicate one. ON CONFLICT DO NOTHING and check-and-skip are the same hazard: they convert a key defect into silence.

The answer shipped is conditional skip: a key hit is a no-op only when the incoming hop re-derives the stored row's hash, and a divergent hit is reported rather than committed away. That is a real narrowing, and it is deliberately not sold as more than it is — a content check cannot see a collision between two content-IDENTICAL events, which is exactly why the key must carry every dimension. Neither half is sufficient alone. Tracked: #2398.


5. Layered-rejection shadowing — the inner layer is never reached​

The predicate (use this one). A rule is shadowed iff the outer layer rejects a superset of what the rule rejects — rejection-set containment. It is not "the rule duplicates an outer constraint"; that framing flags nearly everything and is useless. Strictly-stronger rules stay reachable and pinned — the 200-leg ceiling, computed cross-field relations JSON Schema cannot express, and so on.

Detection is uniform across all three forms: delete the inner rule, run the suite. Green ⇒ shadowed.

Form A — schema → Go (#2392, from PR #2388's 34-mutation sweep)​

The semantic cadence floor interval < minWaitPollIntervalSeconds (#2255) could be deleted with the suite staying green. Its only test entered through ValidateDefinition, where the JSON schema's "minimum": 30 rejects first. Worse than untested: validateSemantics is private, has exactly one caller, is schema-first and returns on failure — the rule was structurally unreachable in production. Since fixed: the cadence floor is now pinned below the schema by TestValidateWait_CadenceFloorIsPinnedBelowTheSchema (internal/workflow/validator_wait_observation_test.go:463), which calls validateWait directly rather than entering through ValidateDefinition.

The population is unmeasured, and that is the honest statement. #2392's sweep — ~40 minLength/minimum/maximum constraints plus 12 enums — has not been run. How many of them shadow a Go rule is unknown. Not "some", not "several": unknown until someone runs the deletion test against each.

Form B — Go gate → SQL predicate (PR #2393, E-4b)​

Three SQL org_id predicates were unreachable because a Go gate (requireVisibleRun) rejects first. Same containment shape, different medium. Retesting below the gate pinned them — and the artifact case needed two orgs sharing a run_id to construct an input the gate admits but the SQL must reject. The test had to be built to reach past the shadowing layer.

A related-but-distinct find from the same PR: a gate satisfied by the gate below it — two sequential gates where the second's rejection makes the first unfalsifiable.

Form C — mutual shadowing (PR #2393 review 4833939651)​

Same layer, same statement, two redundant conjuncts covering each other: an explicit org_id = $2 predicate and the current_setting('app.org_id') RLS guard. Mutating either one alone leaves the suite green. Deleting both is caught.

The fix is different in kind: you cannot pin these by testing below a layer, because there is no layer below. It needs a case where the two conjuncts disagree — here, param-org ≠ tx-GUC-org, which one rejects and the other admits. A single such test kills both mutants.

For any redundant pair of conditions guarding the same thing, a test must exist where the two disagree. Otherwise defence-in-depth is unverifiable by construction — and silently becomes defence-in-one the first time someone simplifies.

The direction that matters most​

The dangerous move is the outer layer being relaxed, not the inner rule deleted. A test asserting only "validation failed" cannot say which layer rejected, so it pins neither. Loosen the schema and the Go rule silently becomes the sole untested gate; delete the Go rule and the schema silently becomes the sole untested gate. Audit the pair, not either half — and assert which layer rejected.

Two technique notes worth carrying: assert on error identity, not a substring probe (#2388's hardened assertion killed an inlined look-alike that satisfied a substring — and the architect's own proposed seam would have passed that mutant); and make a seam t.Fatal if the case under test is already admitted, so a later slice cannot silently render the test vacuous.


6. Routing-edge descent — N hand-maintained copies of one set​

Nine escape instances across four PRs. Every one caught by human review, none by CI. Two reached shipped main (#2270, escalated to P1 — both since closed by PR #2274, merged 2026-07-28; stepExit is on main today. The two below are stated as they stood when #2270 was filed):

  • P4 — user_input.on_timeout_next bypasses the parallel join: a step inside a parallel branch escapes to a post-join successor without joining.
  • P5 — the same edge bypasses the loop exit: escapes the loop region without going through the loop's exit path.

These are containment/correctness bugs, not hop under-counts: a governed run could route around a join or a loop boundary. Probes P3b/P3e reproduced the residual region escape with plain agent_action / approval_gate — no WR2 node type needed. The gap was generic and pre-existing.

The actual defect​

Five independent whole-graph walks each hand-maintained their own edge enumeration: cycleEdgeKeys · hopCalc · validateRegionScopes · regionSuccessors · loopBodyTails.

Nothing forced them to agree, so the #2193 "descend every routing edge" fence was a convention, not a structure: a new node type got descended by whichever passes its author happened to think of. PR #2271 descended 3 of 5 and looked complete.

The clearest evidence that a convention cannot hold this: loopBodyTails was found blind twice, independently, in the same week, by two different agents, on two different node types (#2269 wait, #2271 question) — and in both cases only while chasing a different reported defect.

What actually closed it (PR #2274)​

One vocabulary — type stepExit struct{ Key, Target string; CarriesEnvelope bool } — with three accessors (stepExits/stepExitTargets, stepContainmentTargets, stepRegionTargets), every walk consuming those and nothing else. Routing, containment and envelope-bearing stay three separate sets by explicit architect instruction.

Then the part that makes instance #10 impossible — two CI guards:

  • TestCIGuard_NoHandRolledCycleEdgeKeysWalks — AST-parses the package and fails when any function outside a named, documented allow-list ranges over cycleEdgeKeys. The allow-list is also checked for staleness in reverse.
  • TestCIGuard_EverySuccessorBearingStepFieldIsReachable — reflects over dsl.Step, detects every field whose JSON name looks like a successor (next, *_next, on_*, body, branches — deliberately over-broad), and fails unless it is either declared an edge and actually resolvable through stepRegionTargets, or recorded as a policy word with a reason.

Mutation M1 — dropping the single line case "user_input": add("on_timeout_next") — reddened six tests. That is the proof the centralisation holds: one declaration, every consumer inherits.

The trap inside the fix​

on_timeout on a user_input is a policy word (fail|default|route), not an edge. On a wait, the identically-named field is a branch name and is an edge. Admitting it universally would make every walk chase a step named "route". Hence the guard's second arm: a successor-looking field may be excluded, but only by being recorded as a policy word with a reason — never by silence.


7. Skip-reports-as-green — a required check satisfied by not running​

A path-filtered CI job that does not fire reports success, not "skipped", at the roll-up level. A required status check is therefore satisfied by a job that never ran.

The milestone-#28 instance (#2218, since fixed — see below): the WR2-3 sandbox/workspace security-assertion suite runs through db.yml's Migrate-Up/Migrate-Down integration lane, gated on the internal/workflow/** + cmd/agent-orchestrator/** path filters. A dependency-only change — e.g. to internal/mcp/egress — that could weaken the sandbox's egress/isolation guarantees will not re-fire the suite, because it does not touch the filtered paths. The PR went green with the security assertions unexecuted.

Fixed: PR #2431 moved the two packages into security-unit-tests.yml, which triggers on every PR with no top-level paths: filter, so the suite now runs regardless of which subtree moved. #2218 closed 2026-08-03.

The class is not measured elsewhere. No audit of the repo's other path-filtered jobs has been run, so how many other invariant-bearing suites sit behind a changes filter is unknown — which is the reason the check below is stated as a question to ask, not as a count of known instances.

Two consequences, and the second is the one that costs hours:

  1. Gating: an invariant/security suite must gate path-independently of the change surface. Path filters are a cost optimisation for ordinary lanes; on a suite whose job is to prove an invariant, they are a hole.
  2. Diagnosis: comparing a skipped run against a real run is not a comparison. Read the per-job conclusion (skipped vs success), never the roll-up, before concluding "the same suite passed on the other branch".

8. Version gates unpinnable — the oracle excludes the thing you want to pin​

No replay fixture we have, or could record, can falsify the deletion of a workflow.GetVersion call. Deleting a gate leaves the entire 32-fixture replay corpus green. All 13 gates in the interpreter rested on reasoning alone (#2312).

This is deliberate SDK behaviour, not a corpus gap. In go.temporal.io/sdk@v1.46.0: skipDeterministicCheckForCommand excludes a RECORD_MARKER command named versionMarkerName; skipDeterministicCheckForEvent excludes the matching MARKER_RECORDED event; and incrementNextCommandEventIDIfVersionMarker skips the marker and its paired UpsertWorkflowSearchAttributes event when matching command positions. The exclusion is what lets you delete a GetVersion call once old runs drain. The side effect is that the marker is invisible to our only determinism oracle.

Surfaced because a survived mutation was reported rather than quietly dropped (on #2258 / PR #2310). That reporting discipline is the reason this class is in the retro at all.

What actually pins a gate (#2335)​

  • Tier 1 — AST gate registry. internal/temporalwf/version_gates_test.go enumerates every workflow.GetVersion call with its change-id, const, maxSupported, and its exact lexical scope path, e.g. func dispatchStep > switch(step.Type) > case(dsl.StepWait) > if(step.WaitMode == waitModeSignal && GETVERSION >= 1). The scope path is what catches mis-scoping rather than mere deletion: dropping a guard predicate (widening) or adding one (narrowing) changes the fingerprint, while a cosmetic reformat does not. versionContractCheck has 4 sites and the comparison is a multiset, so deleting 1 of 4 is red.
  • Tier 2 — frozen pre-gate fixtures, recorded against a binary whose DefaultVersion arm is live, for the gates where a live in-flight run could hold the pre-gate shape. The generator asserts the pre-gate property (no marker, no post-gate command) so running it against patched code fails loudly instead of quietly writing a useless history. Hand-editing an existing fixture to strip its marker is not a shortcut — the replayer fails with missing history events, expectedNextEventID=…, because event ids must stay contiguous and are cross-referenced throughout.
  • Tier 3 — the rule that protects Tier 2. Bulk fixture regeneration is a silent gate-unpinning operation. Every *_pregate_* fixture is excluded from every regenerate path and frozen at _v0 by definition.

Three findings nobody expected​

  • Two gates were already pinned by accident and nobody knew. wf-run-log-narration goes red on 11 pre-#1987 histories; wf16-cancel-voids-pending-approval on cancel_v1.json. Nothing recorded it — and a routine regeneration of linear_dag_v1.json or cancel_v1.json would have silently destroyed both.
  • A third category exists: reachable but not falsifiable. The replayer compares the command stream, not payloads or timer durations. lbe-b3-user-input-asked-by only changes an activity's input payload, so it replays green when deleted; the const's own stated rationale was wrong about why. Recorded as a category rather than pretended away. Likewise gate_approve_v1/gate_reject_v1 do not pin wf-gate-repoll — a 15s poll leg matches a 24h wait command-for-command, and the signal resolves before any re-poll command is emitted. The timer had to elapse.
  • The fourth evasion (#2339 → #2342): one token. if GetVersion(...) < 1 { old; return } was voided by deleting the return — a return is not a call, so the callee-set key stayed byte-identical and Tier 1 passed. The fix renders each arm as callees + terminator (return/break/continue/goto/fallthrough/panic, label included — break:legs ≠ break); 11 of 17 registry entries gained +return. Terminators that are calls (os.Exit, an always-panicking helper) are deliberately not enumerated — they already appear in the callee set, so swapping one in changes the key twice over.
  • The author then defeated his own fix and closed it in the same PR. The same gate written as a switch case renders gated=[] other=[], because arm extraction only reads if-conditions — so a gate authored in that shape is born with no arm content and every edit inside it is invisible. Converting an existing gate is caught (the ancestor path changes); a new one would not be. Tier 1 now rejects the shape outright with an actionable message. 52 mutations, 0 unexpected.
  • Residual, recorded not closed: an arm's terminator is pinned but not its operands — return ok → return false keeps the key green (measured). The gate is not voided (the arm still does not fall through) but its result changes.

The generalisation​

A green suite is not evidence that a construct exists, if the oracle deliberately excludes that construct. When it does, no amount of that oracle will pin it — you need a second oracle at a different level (here: structural/AST over source, beside behavioural over history).


The observation that generalises​

Six of the eight classes share one root: a set that had to be enumerated exhaustively was enumerated by hand, per site, by whoever was editing.

ClassThe setCopies
3 / 4dimensions a key must carryone per key-minting site
5layers that can rejectimplicit, never written down
6routing edgesfive walks
8version gatesthirteen, pinned by reasoning

In every case the failure was invisible to review because each site was internally consistent — a hand-maintained list audited against itself always agrees with itself. The only things that ended a recurrence were convergence to one declaration plus a CI guard that fails when a member is added without registration (class 6), or a second oracle at a different level (class 8). Prose rules did not: classes 3 and 6 each recurred after the rule had been stated and acknowledged.

The clustering corollary, which is well evidenced for classes 2, 3, 5 and 6 (and not for class 1 — see the correction there): when you find one instance, the sibling instances already exist. #2382 fixed one input-independent fake and three siblings sat in the same file. #2270 fixed one routing edge and eight more instances followed. Budget the sweep at discovery time, not after the second recurrence.


Where the rules live​

Deliberately split so no rule has two homes:

SurfaceHoldsPath
This retrothe evidence — what happened, in which issue, what measurement settled itdocs/retrospectives/2026-08-03-wr2-failure-classes.md
Lessonsthe author-side construction rules — what to do while writing the code or the testtasks/lessons.md
Reviewer checklistthe review-time questions — what to ask of a diff, and what evidence refutes the concern.claude/agents/principal-architect.md, beside the SHELFWARE CHECK

Each surface links to the others rather than restating them. If a rule needs changing, change it in its one home.

Refs​

#2380 (this write-up) · #2346 · #2384 · #2398 · #2396 · #2397 · #2392 · #2388 · #2393 · #2270 · #2274 · #2271 · #2269 · #2218 · #2312 · #2335 · #2339 · #2342 · PR #2344 (closed) · LLD #2253 · HLD #2223 · PRD #2081 §4.4 · Tracker #2103