Skip to main content

KT — Workflow Runtime v2

Audience: an engineer picking up the workflow engine after v2 (PRD #2081, milestone #28, closed 2026-08-12). Goal: understand what v2 added, know precisely what is live vs dormant vs flag-off, and internalise the one failure class that produced nearly every serious defect in the programme.

Companion doc: Workflow Orchestration is the hub — the WorkflowService API, the execution substrate, run projection and the governance choke-point. This doc is the v2 spoke. Read the hub first if you are new to workflows entirely.

Revision: all claims verified against 70853560 — the revision currently deployed on beta-dev.

Status legend​

🟢 LIVE · 🟡 STUBBED · 🔵 DEV-ONLY / default-off · 🟠 PENDING, PARKED or NOT DEPLOYED


0. TL;DR — the mental model​

Before v2, a workflow was essentially a straight line of agent steps with an approval gate. v2 makes the graph expressive and the runtime durable across time and people:

  • A step declares what it emits; its successor declares what it expects. Mismatches are rejected when you save the workflow, not at 3am.
  • A workflow can loop with a hard iteration cap, and fan out into concurrent branches with join policies.
  • A step can run code in a sandbox whose network is fail-closed by topology, not by an in-process allowlist.
  • A run can park for hours or days waiting on the outside world — by signal, or by governed polling — and resume durably across worker restarts.
  • A run can stop and ask a human a typed question and route on the answer.

What v2 does not do: dynamic decomposition. The graph is fixed at author time. Child workflows (WR2-6) were deliberately carved to v2.1, so there is no data-driven fan-out and no parent/child run lineage.


1. Architecture & control flow​

Walkthrough. 0 A definition is saved and the validator runs its passes in order — region scope, cycle detection, hop proof, contract composition. 0a Anything that fails is rejected at save; this is the primary defence. 1 An accepted definition boots the interpreter with a shared hop budget. 2 A loop re-enters via its back-edge, evaluating until over the body-tail envelope. 3 A park suspends the run durably. 4 Each poll leg derives an idempotency key carrying run, step, lap and leg. 5 Only admitted (tool, action) pairs reach the invoker. 6 code_exec runs in a sandbox with no default route. 7 Every hop is audited and costed.


2. The five areas​

WR2-1 — Per-step I/O contracts 🟢 LIVE​

AspectReality
ShapeFlat FieldSpec, not JSON-Schema — {type, required, description}, type ∈ string|number|boolean|object|array|artifact (dsl/contracts.go:35-48)
Save-time checkValidateContractBlock + SupersetViolations — producer emits ⊇ consumer required expects, matched on name and type (contracts.go:62,127)
Per-edge, not unionEach incoming edge is checked independently; a predecessor with no emits fails every required field (validator_contracts.go:655-682)
Runtime enforcementenforceExpects / enforceEmits / enforceToolEmits (interpreter/contracts.go:86,120,170)
BreachErrTypeContractBreach, non-retryable (api/activities.go:133) — routes via escalateAcceptance honouring on_breach / MaxRedos
Determinism gatewr21-per-step-io-contracts (interpreter.go:979)
Where allowedexpects+emits on agent_action, tool_action; emits-only on parallel, wait, question (validator_contracts.go:37-40)

The per-edge rule matters and is easy to get wrong: a union-of-predecessors check is unsound, because the engine delivers exactly one predecessor's envelope at run time. A field present in some predecessor is not a field the consumer will receive.

Known doc-vs-code defect

validator_contracts.go:13-16 still describes the check as "the union of all direct-predecessor emits". The code at :659-681 is per-edge. The comment is wrong, not the code.

WR2-2 — Bounded loops + parallel fan-out 🟢 LIVE​

AspectReality
MaxIterationsRequired, 1..25 (validator.go:33, validator_loop.go:43-52)
until semanticsDo-while. First entry dispatches the body unconditionally; until is evaluated only on back-edge re-entry, over the body-tail envelope, inside the EvalBranch activity (steps.go:621-651)
Exhaustionfail (default, fail-closed) · escalate (human gate) · route → on_exhausted_next (steps.go:654-683)
Branches2..16 (maxFanOut, validator.go:39)
Join policiesall, any, n_of_m (+n), first_success (parallel.go:43-48)
Loser handlingdrain only — cancel_hard deferred. drainRest blocks on each branch so in-flight cost and audit still accrue (parallel.go:246)
OrderingFutures consumed in sorted order, never a Selector — completion order is non-deterministic and would break replay (parallel.go:364,233)
Hop budgetmaxHopBudget = 200, one shared counter. Loop = mi × (1 + bodyHops) + max(next, on_exhausted_next); parallel = 1 + Σ(branches) + next, additive (validator_loop.go:185-345)
Gateswr22-bounded-loop, wr22-parallel-fanout

WR2-3 — Run workspace + governed code execution 🔵/🟠 BUILT, NOT ENABLED​

This is the area where the delivered code and the deployed reality differ most. Read this section before assuming code_exec works.

AspectReality
Node kind?No. code_exec is not a node kind — it is a high-privilege tool_action action with a catalog ClearanceFloor enforced at save (validator.go:1150-1162)
SeamSandboxRunner.Run (sandbox/runner.go:145); sole impl gvisorRunner (:192)
Default substratedockerLauncher — SANDBOX_SUBSTRATE defaults to "docker" (config.go:351) 🔵
K8s pathK8sLauncher exists but is dormant; the GKE pod-exec transport is a loud fail-closed stub (ErrK8sSubstratePending, k8slauncher.go:120-132) 🟠
gVisor enforcementRuntime defaults to runc (degraded); enforcement default off (config.go:348-349) 🔵
On beta-devOFF. No WF_CODE_EXEC_ENABLED / SANDBOX_* in the container env; default false (tool_catalog.go:222-224), so no provider is constructed (main.go:1225) and GetToolCatalog does not advertise it. No sandbox or proxy containers run. 🟠

Egress — the security claim, stated precisely. The design is fail-closed by topology rather than by an in-process allowlist, because a Go RoundTripper allowlist cannot bind raw sockets and so cannot constrain a subprocess:

ControlEvidenceStatus
Sandbox netns has no default routedocker-compose.sandbox.dev.yml:18-20 (internal: true), pinned at dockerlauncher.go:324-326🟢
Blank network ⇒ refuse to launchdockerlauncher.go:88-90🟢
Forced proxy is default-deny; allowlist consumed as datadev/sandbox/squid.conf:32-36🟢
SSRF floor denied firstsquid.conf:13-18🟢
Empty grant ⇒ no HTTP_PROXY at alldockerlauncher.go:349-353🟢
DeriveEgressGrant fail-closed on nil allowlist / lookup error / IP literalpolicy.go:110-138🟢
The per-run grant is not enforced end-to-end

The derived per-run egress grant is never written to the proxy. No Go code touches allowlist.txt — it is mounted statically from compose (dev) or a ConfigMap that ships empty in prod (egress-proxy-configmap.yaml:84). The revocation obligation is documented-open at policy.go:13-20.

Net effect: the grant is intent and audit only; the enforced posture is deny-all-except-the-static-list. That fails closed, so it is not a vulnerability — but the advertised per-run allow does not work, and anyone enabling code_exec for real traffic must close this first. 🟠

Workspace. Migration 177 creates workflow_run_workspace and workflow_artifact with FORCE RLS. Blobs are content-addressed (content_sha256 is identity) under orgs/{org}/runs/{run}/{sha} — so dedup happens within a run only, never across tenants or runs (blobstore.go:75). Quota 1 GiB/run, 30-day retention (store.go:39-42). The artifact FieldSpec is always a ref, never bytes (dsl/inputs.go:45-51).

Two deployment gaps: the reaper binary is not deployed anywhere (zero compose/k8s references to cmd/workspace-reaper) 🟠, and the orchestrator wires NewMemBlobStore() — an in-memory placeholder (main.go:1279). Note that call sits inside the dormant CodeExecEnabled branch, so it is not reached at all while the flag is off; it becomes a real gap the moment you enable code_exec 🔵.

Resource bounds (docker path): CPU, memory and pids are enforced at container level via --cpus / --memory / --pids-limit 🟢. Disk is a client-side 200ms poll plus docker kill — enforced but overshoot-prone; kernel tmpfs sizeLimit exists only on the dormant k8s path 🟡. Wall-clock and output-byte caps are enforced 🟢. The clamp can never widen the defaults (policy.go:53-84).

Cost. cpuSec × rate + wallSec × rate (placeholder rates, sandbox/cost.go:20-37), joined on iteration so the earlier N× loop over-count is fixed (sandbox_cost_meter.go:117-127, migration 180). Caveat: dockerLauncher never populates CPUSeconds, so dev cost is wall-clock only 🟡.

WR2-4 — Wait for external event 🟢 LIVE​

Two modes.

Signal modePoll mode ("Option E")
EntryrunWaitSignal (wait.go:267)runWaitPoll (wait.go:359)
SourceDeliverWorkflowEvent → st.events[key], delete-on-consumeGoverned ActInvokeTool returning an observation envelope with monotonic leg / observed_at_unix / elapsed_seconds (wait.go:492)
DedupedeliveryDedupeKey(event_key, delivery_id), durable half in migration 188Idempotency key run:step@lap#leg
Cadence floor—30s (validator_wait.go:63)
Ceiling—5 consecutive failures → fails closed (wait.go:212,519-538)
Denial—Fails closed immediately — entry Govern deny at :391, per-leg isFatalObservationErr at :476. A transport error is a non-satisfying leg and polling continues

The observation envelope is non-constant by construction — it always carries a monotonic leg counter — which is what makes a poll surface falsifiable. Governance is split: a coarse ActGovern runs once before leg 1; per-leg tier checks and credentials happen inside InvokeTool.

WR2-5 — Structured human question gate 🟢 LIVE​

runQuestion (question.go:146): open the gate → park → await → unpark → schema-check in an activity → build envelope → N-way route on options[].next via a keyed map read, with an unknown key failing closed (:219-224). Answers arrive on a dedicated SignalQuestionAnswer into a third map (st.questionAnswers) — deliberately never mixed with decisions or answers. Timeout policies: fail / default / route.

WR2-6 — Child workflows 🟠 CARVED TO v2.1​

Out of scope for v2 by founder decision. The consequence for a reader: the graph is fixed at author time. No data-driven fan-out (parallel.branches is a static named list), no runtime decomposition, no parent/child run lineage.

Bonus: the switch node 🟢 LIVE (post-milestone)​

Added after the milestone closed (babe1f09, gate core2538-switch-classifier-router). A conditional is a zero-cost 2-way CEL boolean; a switch spends one classifier model turn and takes N labelled edges, running the full govern→audit→execute→accrue spine. default_next is required and rejected at save if missing — no-match, below-min_confidence and unparseable all route there. Crucially, a govern DENY or guardrail block is not laundered into the default; it is a truthful run failure (switch.go:32-36).


3. The identity-key discipline — read this section twice​

Nearly every serious defect in this programme was one failure class: a construct that repeats along a dimension, whose identity key omits that dimension.

The dimensions in play are step, lap, leg, attempt, ordinal and concurrent-park. Here is the current state of every identity key:

Status cells are a snapshot of the pinned revision (see the footer); the mechanism claims below are maintained continuously and corrected in place when they turn out to be wrong.

KeyDerivationCarriesComplete?
Audit hophop:{step}:{event}:{lap}[:{ord}] (events.go:680)step, lap, leg, attempt, ordinal🟢 since #2398
Poll idempotencyrun:step@lap#leg — waitObservationIdempotencyKey (events.go:175), anchor wait.go:435-437step, lap, leg🟢 since #2397
tool_action idempotencyrun:step@lap — lapToolIdempotencyKey (events.go:155)step, lap🟢 since #2397
Step-execution anchorcoordinator_task_steps(task_id, step_name, iteration) (migration 195)step, lap🟢 since #2442
Parked-approval trackingrunState.pendingApprovals []string — ordered set (interpreter.go:229)concurrent-park🟢 since #2381
Narration event idnarrate:{step}:{tag}:{lap} (narration.go:63,139)step, lap — no ordinal🟠 #2448 — fix approved in PR #2588
Approval-gate dedupsha256(task_id + ":" + step_name) (approvalgate/gate.go:264)step only🟠 #2445 — hardening in PR #2591; see below
Two of seven keys are still short a dimension

The exemplars are fixed; the population sweep is not. The retrospective says so itself (2026-08-03-wr2-failure-classes.md:53-54 — "Population still open", "Population UNMEASURED"). Do not read the four closures below as "the class is done."

#2448 is the live one. agentAskOpenedNarration passes a constant literal tag (agent_ask.go:362), so every ask in a multi-ask turn on one lap derives the same event id. AppendNarration is uuidv5 + ON CONFLICT DO NOTHING (narration.go:84-87), so asks 2..N are silently dropped from the Logs tab. Textbook silent-drop dedupe.

What makes it instructive: the sibling derivations in the same file — askStepName(name, ord) at :343, askEvent(kind, ord) at :352 — do carry the ordinal, so the audit and row paths are complete and only narration is short. A latent secondary: askStepName conditionally elides the ordinal at ord <= 1, which is exactly the positional ambiguity the hop-key shape invariant forbids (events.go:443-448). The generalisable rule the codebase settled on is that a dimension segment is emitted always — @0 outside a loop — never elided at its zero value.

The four that bit us, all now closed:

IssueMechanism
#2398hopKey did not take lap; measured 10 of 19 hops silently dropped on a 3-lap loop. Fixed by making lap a required positional parameter
#2442coordinator_task_steps was lap-blind and anchor-check-first, so a loop body served lap 1's output with zero LLM execution
#2381runState.pendingApprovalID was a single shared slot, so a cancel voided at most one concurrently-parked human gate; the rest leaked as pending on a clearance-gated path
#2397The idempotency anchor carried the leg but not the lap. The poll half was fenced behind an unshipped node type — but the tool_action half was LIVE: create_issue_comment inside a loop body found lap 1's marker on lap 2 and returned lap 1's comment, silently. Closed 2026-08-11 (PR #2558) by lap-scoping both derivations under one change-id, core2397-idempotency-key-lap

#2445 — the one that looks open and is not what its title says. The approval-gate dedup key at gate.go:264 genuinely is lap-blind. It does not produce a bypass, because the dedup index is partial:

CREATE UNIQUE INDEX IF NOT EXISTS ix_gov_approvals_dedup
ON governance_approvals(org_id, session_id, tool_name, args_sha256)
WHERE status = 'pending'; -- 038_approval_chain_engine.up.sql:36-38

Lap 1's row leaves the index the moment it is approved, so lap 2 inserts fresh and the human is asked again. It is P3 rather than P1 because the safe behaviour is incidental: it is a property of a WHERE clause in a migration that no caller of approvalgate can see.

Correction 2026-08-12 — what widening the index actually does

An earlier revision of this doc (and of #2445, and of a PjM comment on it) said that widening the index "silently recreates a real bypass." That is wrong on mechanism, and the true failure mode is better news. Measured twice — by the implementing agent and independently reproduced by the architect on a separate migrated scratch DB — with ix_gov_approvals_dedup recreated TOTAL:

Lap 2's INSERT collides (23505) with lap 1's decided row. pgStore.InsertPending catches the unique violation and falls through to a recovery SELECT — which carries the same AND status = 'pending' filter (store.go:288-294). It therefore matches nothing, and the activity returns:

approval/store: dedup select: no rows in result set

So widening the index gives you a stuck gate, loudly — a retrying activity and a run parked forever at that node — not a silent auto-approval on lap 1's verdict. Fail-closed, not fail-open.

The conclusion is unchanged and still load-bearing: do not widen the index. A run that can never advance is still an outage, and the reasoning that made the predicate look removable ("this dedup helper never fires") is still wrong. Only the failure mode was mis-stated.

Read this before you "fix" the re-asking

The same partial predicate makes AlreadyDecided dead code: an approval gate inside a loop re-asks the human on every lap, and nothing short-circuits it. That looks exactly like a bug — a reasonable engineer sees a redundant prompt, finds a dedup helper that never fires, and widens the index or drops the WHERE status = 'pending' clause to make it work.

Do not make that change. The re-asking is not a defect; it is the safety property wearing the costume of one — a human who approved lap 1 has not approved lap 2.

Be precise about what the change costs, because an overstated warning is one a reader eventually discovers is false and then stops trusting. Dropping the predicate does not hand lap 2 lap 1's approval; the recovery SELECT filters on status = 'pending' too, finds nothing, and the open fails with approval/store: dedup select: no rows in result set (see the correction note above). You get a run wedged at the gate rather than one that advanced without asking. Fail-closed — and still an outage.

Two things follow, and they are not the same thing:

  • To harden this (issue #2445, PR #2591): make the key lap-aware and keep the predicate partial. This is a no-op on today's behaviour — a decided lap-1 row is already out of the index, so there is nothing to collide with. Its entire value is that the safety stops being incidental: with the lap in the anchor, laps never contend for one index entry however the predicate is later spelled. (Demonstrated rather than asserted: with the index recreated TOTAL, the two-laps-two-decisions test still passes; it is the deliberately lap-blind control that reds.) The derivation was hand-copied at four sites in gate.go (:200, :264, :376, :594) — converged onto one dedupAnchor(taskID, stepName, lap) the way PR #2441 converged hopKey. Note the issue enumerated four sites and there were five: acceptanceGateInput (escalateAcceptance) was missed, which is the site the issue's own impact section described.
  • To remove the repeated prompt: no key change can do it. Lap 2 inserts a fresh pending row and reads that row back (gate.go:291), so it is pending whatever the key says. Suppressing the prompt means honouring a decided row across laps — a different query — which widens authorization from "approve this lap" to "approve once, applies to every lap". That is a governance decision requiring founder sign-off, not a refactor.
The lesson generalises beyond keys

This is the same shape as layered rejection: schema rejects before Go does, a gate rejects before SQL does. It can hide a real bug and it can manufacture a phantom one. Reasoning about a key without reading the constraint it is checked against gets you the wrong answer in both directions.

The test discipline that catches it: count, never presence​

ON CONFLICT DO NOTHING and check-and-skip turn a missing identity dimension into invisible loss — the write "succeeds" and a row is gone. A test asserting "lap 2's output exists" passes on the buggy code, because lap 1's row is still there.

So: assert cardinality AND exact members. [1 2 3] and [1 1 1] have the same length, and only one of them is a polling wait.

Three concrete patterns, each worth copying:

PatternAnchorWhy it works
Registry-derived expectationsgoldenflow/provisioner_pollgov_integration_test.go:249The expectation is derived from ObservationSafeActions itself, so seed⊉registry and seed⊋registry both redden. A wildcard target='*' would satisfy a bare count — hence the explicit shortcut probe at :279
Fake purity, non-interference leginterpreter/human_gate_fake_purity_test.go:376-384A fake keyed off a global counter varies on every call and so passes a naive sensitivity check — yet is still input-independent in the way that matters. Only comparing a,a,b against a,b,a on the a entries separates them
Mutation-verified pinswait_composition_golden_e2e_test.go:95 (requireParkedConcurrently)Every outcome a parallel fan-out reaches is also reachable sequentially. Making parallel.go sequential must turn the assertion red — a diverging fixture is not an assertion. Includes a vacuity fence at :129-136

4. The observation-admission fence 🟢 LIVE​

Poll mode cannot call arbitrary tools. Exactly four (tool, action) pairs are admitted (observation_registry.go:120):

ToolAction
platformread_run_state
platformread_artifact_manifest
githubget_check_runs
githubget_pull_request

The bar, quoted from observation_registry.go:19-22: "READ-ONLY — it cannot change the observed system, ever; CHEAP — it is affordable at the leg ceiling; IDEMPOTENT — repeating it changes nothing but the answer." Membership is a review-gated declaration, not a derivation (:17).

Enforcement is create-time fail-closed: validator_wait.go:321 rejects an unadmitted pair at save, and the error lists the admitted set. isObservationSafe is an O(1) lookup built at init from the same slice, so the two cannot disagree.

GetToolCatalog projects this list verbatim to clients (tool_catalog.go:355-362) so the builder renders a supported-list, not a free-text field. The comment at :347 states the rule plainly: this exposes the existing four, it does not add a fifth.

Corrected in place — the two gaps this section used to list are CLOSED

This paragraph previously said the run-time fence lived in the github invoker and that its read-only/marker-free probe was scoped to tool == "github" (#2433, #2435, both milestone #29). Both landed. enforceObservationFence is now provider-agnostic and at the dispatch chokepoint (observation_fence.go, called once from dispatchToolInvoker.Invoke), gating on workflow.IsObservationSafe rather than on any provider-local mirror — so code_exec, preview and platform are fenced by construction — and the evidence suite runs over every admitted pair. Status cells in this doc are a snapshot of the pinned revision; mechanism claims are maintained continuously and corrected in place, because a warning a reader discovers is false is a warning that costs the whole document its credibility.


The governed git capability (core#2568) 🟢 LIVE · 🔒 ENABLED NOWHERE​

The tool_action catalog grew three github rows on 2026-08-12. The interesting one is the first action in the platform that can change a customer's source of truth.

ActionClassFloorObservation-safe?
get_file_contentsread0No — see below
get_branchread0No — see below
commit_fileswrite4No, and never

Catalogued ≠ authorised, and that is the whole model. Nothing seeds a policy for any of the three. The provisioner's two projections — Spec.ToolActionTargets() (from the embedded definitions' own steps) and ObservationGovernanceTargets() (from ObservationSafeActions) — exclude them, and goldenflow/git_capability_seed_2568_test.go reddens if that ever stops being true. Because tool_action is in FailsClosedOnNoMatch, a run reaching such a node on an ungranted tenant fails GovernanceDenied at its entry Govern. Enabling it is a hand-written operator grant: docs/runbooks/governed-git-write-grant.md.

⚠ The T5 guardrails are not the authorisation for the write. EvaluateArguments is fail-open on no-match; only the entry Govern is fail-closed. That is the design issue's own central caution and the easiest thing here to get wrong.

What holds with no policy at all (cmd/agent-orchestrator/github_commit.go — six fences, in the order they run): the literal repo pin; a three-arm protected-ref deny (static conventional names with no network call, the repository's live default_branch, GitHub's protected flag — all case-folded, because GitHub refs are case-sensitive so Main is not main to a naive ==); a path fence that additionally refuses any case/dot-confusable of .github/ and .git/ (.GitHub, .github., ..github all normalise to the denied root, because on a case-insensitive filesystem such a tree path materialises inside the real directory — and a workflow file there executes with the repository's CI credentials, a strictly larger capability than the commit); required expected_head_sha; and force: false on the ref update so GitHub itself rejects a non-fast-forward. There is no force-push and no way to compose one — the four Git-Data primitives were rejected in favour of one composite precisely so the ref move is not separately authorable.

A CAS failure is ErrTypeToolCASConflict — its own class, deliberately absent from toolOpts' non-retryable list, because the retry is what lets a first attempt whose ref update landed but whose response was lost find its own idempotency trailer and return that commit rather than write a second.

Idempotency is the §3 anchor at a fourth layer, and the marker is a real git TRAILER. The commit message carries an Upsquad-Idempotency-Key: run:step@lap trailer, matched only in the trailer block — the last paragraph, and only when every line of it is a well-formed Token: value trailer (git-interpret-trailers semantics). It is deliberately NOT a scan for the prefix on any line: message is agent-produced, so a line-scan let an agent forge a future lap's marker into the commit BODY (making that lap silently skip its write) or adopt a third party's commit into the audit record (PR #2594 review B1, CRITICAL). A belt-and-braces write-gate additionally refuses an author message that contains the marker prefix at all, and the key itself is fenced to [A-Za-z0-9:@#._-] so it can never span trailer lines. Line-exactness within the block still holds — run:step@1 must not match run:step@10. And because the key is baked into a commit object that outlives the run, changing the derivation later is version-gated, same as #2397.

Why neither read joins the observation registry — the classification decision, recorded because the design asked for one: a 128 KiB file body at the 200-leg ceiling is ~25 MB of Temporal history, failing the "cheap" bar; no PRD #2081 §4.4 case asks to wait on file content; and since #2564 admission would auto-seed an org-wide grant on every provisioned tenant, which is the opposite of this capability's rule. get_branch passes the first two comfortably and is still excluded on the third. Pinned by TestGitCapability_IsNotObservationSafe, which also holds the registry at exactly four.


5. Goal 6 — the north-star SDLC-parity E2E 🟢 LIVE​

The programme's headline claim was not any one construct — it was whether the engine could run our own pipeline. northstar.Definition() (northstar/northstar.go:179) is one dsl.Definition spanning five legs:

design ──▶ build ──▶ ┌ review_loop (max 3) ┐ ──▶ await_ci ──▶ merge_gate ──▶ done
│ author ⇄ reviewer │ (poll) (question)
└ until approved:true ┘

It ran live against real GitHub: 16 real get_check_runs calls 30s apart over 7m41s with CI genuinely progressing (3 pending → 2 → 1 → success on leg 16), 16 audit hops under 16 distinct keys, merge-gate answered {"option_key":"merge"} → merge_pr executed.

Frozen as two SHA-pinned fixtures (northstar_fixture_test.go:79): northstar_sdlc_v1 (32d20f03…, laps 1-3, option merge) and northstar_escalate_v1 (aed54e73…, option hold, escalated).

The CI lane (golden-flow-e2e.yml, job northstar) is deliberately unfiltered — no paths:, no if:, no needs: — because a path-filtered job SKIPs and reports success, which would make an acceptance bar meaningless. It also counts executed PASSes rather than trusting exit colour.

Not a required check

The lane is not in the required roster (8 contexts at this revision). It is covered transitively by the required unconditional Go Unit Tests, but the named acceptance context can be bypassed. Promotion is tracked as #2504, blocked on the non-ASCII arrows in the job name — required contexts match byte-exact.


6. Determinism — 18 gates and a null oracle​

Every behaviour change to workflow code sits behind a named GetVersion gate, registered in version_gates_test.go:213 with three tiers of proof (AST site-key walk, pre-gate fixtures, SHA-256 freeze).

Current ids: wf-run-log-narration, wf16-cancel-voids-pending-approval, wf16-cancel-voids-every-parked-approval, wf-gate-repoll, wf26-conditional-cel-routing, wf30-tool-action-node, wf-acceptance-outcome-verification, wr21-per-step-io-contracts, wr22-bounded-loop, wr22-parallel-fanout, lbe-b1-user-input, aiu-agent-ask-user, lbe-b3-user-input-asked-by, wr24-wait-for-external-event, wr24-waiting-projection, wr25-question-gate, core2538-switch-classifier-router, core2397-idempotency-key-lap.

Two rules a new contributor must internalise:

  1. Never range a map in workflow code. Keyed reads only; ordered collections are slices sorted explicitly (LLD #2124 §6 rule 1; cited at interpreter.go:601, parallel.go:343).
  2. A green replay corpus does not prove a gate exists. The SDK excludes the Version marker from the determinism comparison in both directions, so replay is a null oracle for gate presence (doc.go:44-56). The Tier-1 AST walk is the real guard.

core2397-idempotency-key-lap is the sharpest illustration: it is tier-1 registry only, unpinnable by replay. Both arms emit an identical command stream differing solely in the InvokeTool input payload, which the oracle does not compare — deleting the gate replays green against every fixture in the corpus — 55 of 55 at this revision (note the registry comment at version_gates_test.go:593 still says "54/54"; it predates wait_poll_lap_key_v1.json and is itself stale — count the directory, do not trust the number). The fix is falsified instead by direct assertions: TestWaitPoll_ObservationKeys_AreDistinctPerLap and TestToolAction_IdempotencyKey_IsDistinctPerLap. If you add a gate whose arms differ only in payload, say so in the registry and write the direct test — the corpus will not catch you.


7. Gaps, parked and dev-only — the honest state​

Deployment state first, because it is easy to misread. The engine is default-off but on where it matters: docker-compose.dev.yml:688,1097 default TEMPORAL_ENABLED to false, but the running beta-dev orchestrator has TEMPORAL_ENABLED=true and TEMPORAL_WORKER_ENABLED=true (verified on the live container at 70853560). So v2 workflows genuinely execute on beta-dev. code_exec is the exception — it is off there too.

ItemStatusWhat it means for you
code_exec on beta-dev🟠 OFFFlag default false; no provider constructed, not advertised in the catalog. WR2-3 is built but not reachable
Per-run egress grant🟠 Not enforcedIntent + audit only; enforced posture is deny-all-except-static. Fails closed, but close this before enabling code_exec
Workspace reaper🟠 Not deployedcmd/workspace-reaper has zero compose/k8s references
Blob store in orchestrator🔵 In-memoryNewMemBlobStore() placeholder
K8sLauncher / GKE🟠 DormantLoud fail-closed stub; docker is the default substrate
gVisor enforcement🔵 Off by defaultRuntime defaults to runc (degraded)
#2433 / #2435🟠 Open (ms #29)Fence generalisation; #2435 before the registry widens
#2448🟠 P3, live on main — PR #2588 approvedNarration event id is ordinal-blind — a multi-ask turn logs ONE line; asks 2..N are silently dropped from the Logs tab
#2445🟠 Open, P3 — PR #2591Lap-blind approval-gate key, safe only by the partial index. The issue's original title asserted a P1 bypass; refuted, title corrected 2026-08-12. Mechanism correction (same date): widening the index does not silently recreate a bypass — the recovery SELECT filters status='pending' too, so the open fails closed with dedup select: no rows in result set (a stuck gate, loudly). Measured twice, independently. The predicate is pinned as partial from PR #2591 (TestDedupIndexPredicate_IsDeclaredPartial plus a behavioural decided/pending pair, in the db.yml integration lane)
#2552🟠 OpenFake registry omits VoidApproval; require.Len detects removal but cannot discover addition. An enumeration audited against itself proves nothing — three instances found on 2026-08-12 (this, a grep guard satisfied by a matching comment, and #2445's builder list). Worked antidote: interpreter/gate_lap_builder_set_test.go derives the set from the package AST and checks the property on whatever it finds, so a new site must be correct rather than listed
#2485 / #2487🟠 OpenNo guard against a migration numbered below main's ceiling — golang-migrate silently skips it. Duplicates; collapse on pickup
#2494🟠 Open, P2Comments contradict the code on a load-bearing org-isolation predicate (the reconciler role has rolbypassrls, so RLS is inert there)
#2495🟠 Openwf_reconcile_drift_detected_total has no PrometheusRule — drift is counted, nothing alarms
#2504🟠 OpenPromote the north-star lane to required; needs an ASCII context name first
Sandbox Egress Smoke🟠 Path-filteredWorkflow-level paths: means no check is emitted at all for Go-only changes
Security-Relevant Unit Tests🟠 Not required-race lane with 9 assertions, but not a merge gate
No unreachable-step pass🟠The validator is otherwise exhaustive on graph shape, but unreferenced steps are accepted at save
Parallel join-side contracts🟠 DeferredBranch-tail → post-join composition is not validated at save (validator_contracts.go:456-462); on_breach fail-closed is the backstop. No tracking issue cited in code

8. Source map​

ConcernPath
DSL types, node constantsinternal/temporalwf/dsl/dsl.go, dsl/contracts.go, dsl/inputs.go
Validator (schema, limits, pass order)internal/workflow/validator.go
Contract compositioninternal/workflow/validator_contracts.go
Loop / parallel / switch / wait validationinternal/workflow/validator_{loop,parallel,switch,wait}.go
Observation admissioninternal/workflow/observation_registry.go
Interpreter core, iterState, hop budgetinternal/temporalwf/interpreter/interpreter.go
Loop + step executioninterpreter/steps.go
Parallel fan-outinterpreter/parallel.go
Wait (both modes)interpreter/wait.go
Question gateinterpreter/question.go
Identity keys (hop, idempotency)interpreter/events.go, interpreter/narration.go
Determinism gate registryinternal/temporalwf/version_gates_test.go
North-star definitioninternal/temporalwf/northstar/northstar.go
Sandbox runner / launchersinternal/workflow/sandbox/{runner,dockerlauncher,k8slauncher}.go
Approval-gate dedup keyinternal/coordinator/approvalgate/gate.go
Egress policyinternal/workflow/sandbox/policy.go, dev/sandbox/squid.conf
Workspace store / blobsinternal/workflow/workspace/{store,blobstore,reaper}.go
Migrations (internal/context/store/migrations/)177 (workspace), 178 (governance), 179 (artifact ledger), 180 (marker iteration), 188 (delivery dedupe), 195 (step iteration)

Flags & environment​

FlagDefaultEffect
WF_CODE_EXEC_ENABLEDfalseConstructs the code_exec provider and advertises it in GetToolCatalog
SANDBOX_SUBSTRATEdockerk8s selects the dormant K8sLauncher
SANDBOX_RUNTIMEruncrunsc selects gVisor
(gVisor enforcement)offWhen on, a degraded runtime refuses to launch

Node kinds are not feature-flagged. The GetVersion ids are replay-determinism markers, not toggles — every node kind listed here is live on any workflow whose Enabled row is set.


  • Workflow Orchestration — the hub: API, substrate, projection, governance choke-point
  • CI/CD Supply-Chain Security — privileged Actions jobs, App-key reachability, and which controls survive workflows: write. Read it before touching any workflow that mints a token
  • Temporal Architecture & Licensing
  • docs/retrospectives/2026-08-03-wr2-failure-classes.md — the failure-class retrospective written during the programme
  • PRD #2081 (v1.3) · tracker #2103 · milestone #28 (closed) · milestone #29 (tenant-usable)
  • Tenant-facing surface: upsquad-ai/upsquad-client#726

Reflects 70853560, the revision deployed on beta-dev as of 2026-08-12, plus the core#2568 governed git capability merged after it (see §4). This doc is invalidated by: any change to ObservationSafeActions; enabling WF_CODE_EXEC_ENABLED or moving off the docker substrate; closing #2445; adding a node kind; or adding a mutating tool_action. Verify status claims against code before relying on them — several here are deliberately narrower than the milestone's own summary.