KT — Agent Memory Loop
Audience: an engineer picking up anything under internal/context/memory/, the extraction path, the PII detector, or memory recall.
Goal: understand the loop end-to-end, know what is actually on, and avoid re-creating the defects this milestone spent a month removing.
Status legend: 🟢 LIVE (merged, in the prod code path) · 🟡 STUBBED (placeholder, works but not final) · 🔵 DEV‑ONLY (only wired for dev/local) · 🟠 PENDING/PARKED (designed and/or coded but gated off / blocked).
Delivered as milestone 27 (PRD #2037), closed 2026-08-29 — 38 issues, QA-verified at 17/20 acceptance criteria, PjM + PM signed off on tracker #2063.
0. TL;DR — the mental model
- An agent's session ends → the platform mines it for durable facts → a governance gate decides what is allowed to persist → survivors are embedded → future sessions recall them. That is the whole loop.
- The model is never trusted with governance. It proposes candidates; every decision that keeps a human out of the loop is deterministic, server-side, and written down. This is the single most important idea in the subsystem, and it is the product of two ADRs and one ruling (§3).
- Fail-closed is the default everywhere, structurally. An unknown PII class mutates. An uncalibrated model cannot drop. A missing detector quarantines everything. None of these are conventions — they are zero-values and absent-map-key behaviours, chosen so that forgetting to enrol something is safe.
- Vectors are stamped and never compared across stamps. Cosine geometry is not portable between embedders, so every recall path filters on the model that produced the vector. A mixed corpus that ignored this would mis-rank silently.
- Extraction is OFF in production and stays off until three conditions are met (§8). It is ON in dev.
1. Architecture & control flow
Colour key: blue = runtime/worker · amber = governance gate or decision · green = durable store · purple = model call · red = irreversible/absorbing state · grey = shipped but disabled.
Walkthrough
| # | Step | Source |
|---|---|---|
| 0 | Session reaches a terminal state; a run row is inserted ON CONFLICT DO NOTHING | extraction/enqueue.go:37 |
| 1 | Redis consumer group drains the stream | extraction/consumer.go:163 |
| 2 | Transcript + prompt (carrying the non-derivability rule) go to the extractor | extraction/extractor.go:435 |
| 3 | Envelope parsed strictly — DisallowUnknownFields | extraction/candidate.go:87 |
| 3a | One bad candidate rejects the whole envelope → retry → dead_letter | extraction/consumer.go:237, ledger.go:140 |
| 4–5 | Provenance stamped, then deterministic classification | extraction/provenance.go:153, classify/constraint.go:173 |
| 6 | Redaction seam dispatches by class | redact/detector.go:133 |
| 7 | Tier policy decides mutate-vs-flag | redact/tier.go:252 |
| 8 | Gate escalates on any hit; otherwise the row activates | extraction/gate.go:150 |
| 9–10 | Human verdict, then deferred embed on approval | review/reviewer.go:238 |
| 11 | The two vector recall paths filter on the embedder stamp | mcpsrv/pgbackend.go:224, assembly/memory_loader.go:277 |
2. Component catalog & status
| Component | Path | Status | Notes |
|---|---|---|---|
| Extraction enqueue | extraction/enqueue.go:37 | 🟢 | Gated on MEMORY_EXTRACTION_ENABLED |
| Extraction worker | extraction/consumer.go:163 | 🟢 | Redis consumer group |
| Provenance stamp | extraction/provenance.go:153 | 🟢 | (model, endpoint class, prompt rev); AST-guarded |
| Deterministic classifier | classify/constraint.go:173 | 🟢 | Two arms + context guards |
| Confidence gate | extraction/candidate.go:155 | 🟠 | Ships empty — cannot drop anything |
| Redaction seam | redact/detector.go:133 | 🟢 | REDACT_DETECTOR, code default regex |
| gliner NER hybrid | redact/nerdetector.go:146 | 🟢 | ner in dev; sidecar + canary |
| National-ID checksums | redact/natid.go:174 | 🟢 | Go-side, not the model |
| Tier policy | redact/tier.go:252 | 🟢 | Fail-closed zero value |
| Review gate | extraction/gate.go:150 | 🟢 | Escalate-only, tier-blind by design |
| Approval resolution | review/approval_resolve.go:104 | 🟢 | Voids paired approvals on terminal transitions |
| Sampled audit | memaudit/gate.go:115 | 🟢 | 10% (1000bp) — but see §8, no verdicts yet |
| Embed worker | memory/embed_worker.go:217 | 🟢 | |
| Stamp-scoped vector recall ×2 | see §1 step 11 | 🟢 | The RAG path (retrieval/vector.go) got the same fix but queries context_embeddings, not memory |
| Stamp-exclusion counter | recall/stampcliff.go:71 | 🟢 | Makes the mixed-corpus cliff visible |
| Near-dup collapse | memory/neardup.go:167 | 🟠 | OFF; threshold 0.10, ceiling 0.35 |
| Prompt-Guard classifier | compose :945 | 🟡 | ML_DETECTOR_USE_FAKE_CLASSIFIER=true — upstream repo is gated |
StubExtractor | extraction/extractor.go:59 | 🔵 | Fallback when no API key resolves |
3. The governance model — read this before changing anything
Most of this subsystem's shape comes from rulings, not requirements. Code that looks arbitrary is usually a settled argument. Three documents govern:
ADR-0029 — PII detection stays inside the trust boundary
Detection runs self-hosted, in-cluster; managed cloud DLP was rejected on BYOK grounds (a tenant who chose non-GCP would have had all session-derived content inspected by Google). Consequences that show up in code: the model artifact is baked and digest-pinned, Hub fetches are fatal, and the detector must work air-gapped.
It also established class-tiered mutation: quarantine already contains the risk, so mutating a quarantined row destroys the reviewer's evidence without protecting anything. The trigger was a real incident — a 13-digit Unix timestamp inside a code sample was matched as a phone number and irreversibly destroyed the artifact, because agent_memory has no pre-redaction original.
That trade rests on a premise (D3 (a)) which is a constraint on any future reader, not just a justification: a LOW-tier pending row holds un-redacted content at rest, and this is acceptable only because the reviewer is a human inside the owning org who already has access to the source session. Any surface that widens who can read a pending row — a new listing API, an export, a support tool, a cross-org view — breaks the premise the founder approved, not merely a convention.
ADR-0030 — model drops must be recorded
A model classifier may remove a candidate before write only if the removal is recorded — countable, attributable to a model version, samplable. A drop with no record is forbidden. The harm being prevented is invisibility, not the model having an opinion: a drop filter's false positives are invisible by construction, because nothing counts what was never written.
Two consequences: attribution requires real provenance (hence §1 step 4 existing at all), and a model drop is admissible only on a candidate whose redaction class set is empty — redaction is an admissibility test, not a sanitiser.
⚠️ ADR-0030 is decided-only, not yet structural. Read D4 as a rule you must uphold, not one the code enforces for you.
gate.go:91still runspassesConfidencebeforeRedactContentat:95— so today the "no drop on a redaction-flagged candidate" invariant is held only becauseconfidenceThresholdsships empty. The moment anyone populates that map without reordering those two calls, the ADR is violated in code while every test stays green. See §8 for the real prerequisites.
It also carries a residual the ADR names rather than closes (D5(d)): the rule binds only post-emission drops. An extractor prompt tightened to suppress candidates emits nothing to record, so defeating ADR-0030 in substance requires no code change at all — just a prompt edit. Anyone changing the extraction prompt is making a governance change, whether or not it looks like one.
The #2322 ruling — governance vs precision
The extractor gets no vote on governance. high_stakes and importance are server-derived. confidence was carved out as a precision filter rather than a governance gate — a carve-out ADR-0030 later narrowed rather than overturned.
4. State machine & the review gate
Every transition is CAS'd in updateStatus (review/reviewer.go:537); zero rows affected ⇒ ErrIllegalTransition.
| From → To | Op | Notes |
|---|---|---|
pending → active | Approve | Enqueues the deferred embed |
active → active | Approve | Idempotent, no re-embed |
pending → rejected | Reject | |
rejected → rejected | Reject | Idempotent — this is the repair path |
active|pending → tombstoned | Forget | |
active|pending → superseded | Correct | Chains via prev_memory_hash |
active → superseded | near-dup collapse | Sets superseded_by; CAS on status='active' |
rejected → tombstoned | Forget | ILLEGAL — ErrIllegalTransition |
Terminal: rejected, superseded, tombstoned. Only pending and active have outbound edges.
Every terminal transition — including the idempotent branches — also calls resolveOverride, voiding a still-pending paired approval with reason='operator_override_by_2507'. That idempotent branch is deliberate: re-running the operation is the repair for a stranded approval, which is why no one-off cleanup tool exists.
The gate itself (extraction/gate.go:150) is escalate-only and tier-blind by design: any redaction hit forces pending. It can raise scrutiny, never lower it.
5. Operate it
| # | Task | Tool | Effect |
|---|---|---|---|
| 1 | Review a memory | cmd/tools/memory-review -op approve|reject|forget|correct | Terminal transition + paired-approval resolution + audit row |
| 2 | Re-run the classifier over a backlog | cmd/tools/memory-status-reevaluate -dry-run | Read-only decision diff; live pass needs founder approval |
| 3 | Recover dead-lettered runs | cmd/tools/memory-extraction-replay | Step 4 approval-repair is mandatory or recovered memories are invisible |
| 4 | Repair pending rows with no approval | cmd/tools/memory-approval-repair | Fixes the #2323 stranding class |
| 5 | Backfill embeddings | cmd/tools/memory-backfill | Idempotent; near-dup collapse forced OFF |
| 6 | A/B a prompt change | cmd/tools/memory-eval-ab | Real model calls — deliberately never in CI |
Runbooks live under docs/runbooks/memory-*.md. Known defect: memory-review never exits after printing success (#2660) — a timeout wrapper returns 143 over already-applied work. Verify the DB, not the exit code.
6. Data model & API reference
| Table | Key fields | Source |
|---|---|---|
agent_memory | status, memory_type, content_hash, extractor_version, superseded_by, audited_at/audit_verdict/audit_source | mig 005, 169, 196, 201 |
agent_memory_embeddings | embedding vector(1536), embed_model, HNSW cosine | mig 170:50,75 |
governance_approvals | action_type='memory_review', target=memory_id, reason | mig 173:32 |
agent_audit_log | Hash-chained per session; memory rows are session-less | mig 036, review/audit.go:65 |
context_embeddings | model_version NOT NULL, vector(1536) | mig 004:47 |
Migrations in this arc: 169 (loop columns) · 170 (embeddings) · 171 (run ledger) · 173 (review gate) · 184 (replay) · 189 (importance default) · 191 (sampled audit) · 196 (superseded_by) · 197 (memory_reclassified) · 199 (orphaned-approval repair) · 200 (expiry backfill) · 201 (audited marker).
Surfaces: MCP recall_memory / remember (mcpsrv/tools.go:47,69) · PullContext with findLatest / findBySnapshot, both status='active'-scoped (service.go:517,550) · MLDetectorService.DetectPII (proto/upsquad/ml/v1/detector.proto:34).
status='active' is the same allow-list predicate on all seven recall surfaces — service.go:524,558 (PullContext's two arms), mcpsrv/pgbackend.go:58,218, assembly/memory_loader.go, runtime/session/warmstart.go:176, runtime/planmemory/recaller.go:127. All seven carry it, and an eighth must too — this was a real defect (ADR-0029 D5.2, where PullContext returned exactly the row that was supposed to be withheld).
Only two of the seven are vector searches, and only those carry the embed_model stamp filter (§7 trap 3): mcpsrv/pgbackend.go and assembly/memory_loader.go. The other five retrieve by id, recency or snapshot and have no vector to mis-compare — internal/runtime/ contains zero embed_model references, correctly. A new surface needs the stamp filter iff it does similarity search.
Metrics: 33 emitters under memmetrics, every one with a non-test caller, all label vocabularies bounded and enum-sourced. extractor_version labels must use Provenance.MetricLabel() (bounded), never Stamp().
7. Traps — the section that will save you a day
1. Tiering is fail-closed by zero-value. Do not "helpfully" default a new class.
TierFor is a bare map lookup; an absent key yields TierHigh (mutate). Enrolling a class as LOW is an explicit, reviewed act, and it adds a row to the BuildSummary pin. The tier table as shipped:
| Tier | Classes | Behaviour |
|---|---|---|
| LOW | phone, ip, person, dob, address, detector_unavailable | Flag + quarantine, content left verbatim |
| HIGH (by absence) | secret, ssn, cc, email, natid | Mutate + quarantine |
2. A quantised model detects nothing while reporting healthy.
Measured during selection: the int8 gliner variant returned zero detections, health-green, and faster. That is why the shipped artifact is fp32 and why the health probe asserts a canary detection (a known-PII string must produce hits) rather than liveness — and why the compose healthcheck targets MLDetectorService.DetectPII, not the always-SERVING service key.
3. Never compare vectors across embed_model / model_version.
Cosine geometry is not portable. All three recall paths filter on the live embedder's ModelVersion(). A mixed corpus without this mis-ranks with no error, no metric, nothing red. memory_recall_stamp_excluded_total exists so the exclusion is visible rather than looking like an empty corpus.
4. dead_letter is absorbing.
claimRun drops terminal runs; only an explicit replay CAS re-arms one. A run that dead-letters stays dead until someone runs the replay tool — and the replay runbook's step 4 (approval repair) is mandatory, or the recovered memories exist but are invisible to reviewers.
5. A bad candidate rejects the whole envelope.
ParseEnvelope uses DisallowUnknownFields and fails the entire run on one malformed candidate. This is why the non-derivability rule rides the existing Rationale field rather than a new required field: a new required field would convert one model formatting lapse into total memory loss for that session.
6. Two retention mechanisms, and only one of them works on loop-written memories.
The org-level RTD sweeper (4h tick, agent_memory at 180d default / 30d floor) does reach these rows — they are not immortal. But the per-row expires_at / retention_days columns are written only by PushContext; the extraction, remember and Correct INSERTs omit them, so every loop-written memory has expires_at NULL and is invisible to idx_memory_expires_at. Per-memory_type decay (AML-F16) is genuinely not built.
7. Comments that disagree with the code (correct as of d3d8886e, fix them when you touch them):
| Where | The lie |
|---|---|
redact/detector.go:12-20 | Says the seam has no NER implementation and ner falls back to regex. False since #2155 — only the dlp arm still matches |
extraction/gate.go:1-6 | Reads as three live gates; the confidence gate cannot drop (empty map) |
extraction/extractor.go:4-6 | Calls StubExtractor the default wiring; the composition root prefers GeminiExtractor whenever a key resolves |
8. Code default ≠ deployed default. MEMORY_EXTRACTION_ENABLED is OFF in code, true in compose. Reading only one will mislead you.
EMBEDDING_PROVIDER used to be the worst instance of this — openai in code, vertex in compose, so which third party received a tenant's documents depended on which layer you read (ADR-0029). Both defaults are deleted by LLD #2714 T4 (PRD #996 FR-9 step 1, #2546): unset is now a refusal and a non-zero exit with a message naming the choices, and docker-compose.dev.yml states vertex outright rather than as a ${...:-vertex} fallback. There is no longer a layer to read it off; the only answer is the one in the manifest and the one on the boot line.
8. Gaps, pending, parked — the honest state
| Item | Ref | State |
|---|---|---|
| Prod extraction enablement | ADR-0029 D4.4 | 🟠 OFF until C1–C3 below |
| ↳ C1 live dogfood run under the shipped prompt | — | pending |
| ↳ C2 human-rated precision sample | — | pending |
| ↳ C3 extraction spend metered | #2669 | 🟠 prod-enablement — no llm_usage_events, spend unbillable and uncappable |
| Sampled audit produces 0 decided verdicts in ~4 months | #2654 | 🟠 "samplable but not sampled" — the state ADR-0030 D2 rejects |
| Confidence gate inert (empty thresholds) | candidate.go:155 | 🟠 Do not populate this map casually. It is what currently holds ADR-0030 D4 — see below |
| Near-dup collapse | neardup.go:167 | 🟠 OFF; live active corpus has no true near-duplicates at any safe threshold |
| Prompt-Guard classifier is a fake | compose :945 | 🟡 Upstream repo gated |
| Detector image 9.21 GB (CUDA wheels, CPU-only service) | #2666 | 🟠 Defeats the footprint rationale of the model choice |
Per-memory_type decay (AML-F16) | — | 🟠 Not built; blocks the PRD §10 decay/retention tier claims |
| Tier consolidation (AML-F18) | — | 🟠 Phase 3, not due |
context_shares empty OrgIDColumn | scopes.go:658 | 🟠 Opts out of the generic deleter |
memory-review never exits | #2660 | 🟠 Zero test files in that package |
What enabling the confidence gate actually costs
The single most likely way to break ADR-0030 is to "just add a threshold". Populating confidenceThresholds turns the model into a live drop gate, which the ADR permits only once all of D1–D4 hold. The full prerequisite set, not the label count alone:
| # | Prerequisite | State |
|---|---|---|
| D1 | A drop-record store that retains the dropped content — a counter is not "samplable", and memory_candidates_dropped_total already exists, so counter-only would rename the status quo as compliance | not built |
| D2 | A mandated non-zero sample rate over that store — "samplable but not sampled" is explicitly rejected | not built; and #2654 shows the existing audit stream at 0 verdicts |
| D3 | Real provenance so drops are attributable | ✅ shipped (#2584) |
| D4 | The drop evaluated after redaction, so it can never remove a flagged candidate — gate.go:91 still runs it before :95 | not done |
| — | A classregistry scope for the new retained-content store (class, retention, OrgIDColumn, working Deleter) | not built |
| — | P0.1.3 DPA widening — retained never-promoted content is a new purpose and category, and DSAR/erasure must reach it | not built |
The "≥11 verdicts rendered against the current predicate" figure is the floor-promotion condition for scoring a stamp — a much later and smaller question than switching the gate on.
9. Key files, flags & environment
Flags (code default → dev value)
| Var | Code default | beta-dev |
|---|---|---|
MEMORY_EXTRACTION_ENABLED | OFF | true |
REDACT_DETECTOR | regex | ner |
ML_DETECTOR_ADDRESS | k8s DNS | ml-detector:50060 |
EMBEDDING_PROVIDER | none — refuses to boot (#2714 T4) | vertex (stated in compose, not defaulted) |
EMBEDDING_DIM | 1536 | 1536 |
MEMORY_NEAR_DUP_COLLAPSE_ENABLED | OFF | false |
MEMORY_NEAR_DUP_MAX_DISTANCE | 0.10 (ceiling 0.35, refuses rather than clamps) | unset |
MEMORY_EXTRACTION_TIMEOUT | 30s | 120s |
EMBEDDING_DIMconfigures only the embedder client. Bothvector(1536)literals are unreachable from it — changing it produces a runtime insert error, not a schema change.
Source map — extraction internal/context/memory/extraction/ · classification .../classify/ · redaction + tiering .../redact/ · review .../review/ · recall .../mcpsrv/, internal/context/assembly/, internal/context/retrieval/ · detector sidecar services/ml-detector/ · CLIs cmd/tools/memory-*.
CI lanes — go-test.yml (unconditional, carries the eval floors) · db.yml → Migrate Up + Down (six -tags integration memory steps, per-issue anti-skip needles) · ml-detector-tests.yml (pytest + the real-model corpus floor) · agent-worker-tests.yml. The last two exist because those suites had never run in any workflow.
Related: ADR-0029 · ADR-0030 · runbooks docs/runbooks/memory-*.md · PRD #2037 · tracker #2063.
Doc reflects upsquad-core@d3d8886e (2026-08-29). Invalidated by: prod enablement flipping, #2654 producing verdicts (the confidence gate becomes promotable), #2666 changing the artifact, or any new recall surface — which must carry both the status='active' and embed_model predicates.