Skip to main content

KT — Agent Memory Loop

Audience: an engineer picking up anything under internal/context/memory/, the extraction path, the PII detector, or memory recall. Goal: understand the loop end-to-end, know what is actually on, and avoid re-creating the defects this milestone spent a month removing.

Status legend: 🟢 LIVE (merged, in the prod code path) · 🟡 STUBBED (placeholder, works but not final) · 🔵 DEV‑ONLY (only wired for dev/local) · 🟠 PENDING/PARKED (designed and/or coded but gated off / blocked).

Delivered as milestone 27 (PRD #2037), closed 2026-08-29 — 38 issues, QA-verified at 17/20 acceptance criteria, PjM + PM signed off on tracker #2063.


0. TL;DR — the mental model​

  • An agent's session ends → the platform mines it for durable facts → a governance gate decides what is allowed to persist → survivors are embedded → future sessions recall them. That is the whole loop.
  • The model is never trusted with governance. It proposes candidates; every decision that keeps a human out of the loop is deterministic, server-side, and written down. This is the single most important idea in the subsystem, and it is the product of two ADRs and one ruling (§3).
  • Fail-closed is the default everywhere, structurally. An unknown PII class mutates. An uncalibrated model cannot drop. A missing detector quarantines everything. None of these are conventions — they are zero-values and absent-map-key behaviours, chosen so that forgetting to enrol something is safe.
  • Vectors are stamped and never compared across stamps. Cosine geometry is not portable between embedders, so every recall path filters on the model that produced the vector. A mixed corpus that ignored this would mis-rank silently.
  • Extraction is OFF in production and stays off until three conditions are met (§8). It is ON in dev.

1. Architecture & control flow​

Colour key: blue = runtime/worker · amber = governance gate or decision · green = durable store · purple = model call · red = irreversible/absorbing state · grey = shipped but disabled.

Walkthrough​

#StepSource
0Session reaches a terminal state; a run row is inserted ON CONFLICT DO NOTHINGextraction/enqueue.go:37
1Redis consumer group drains the streamextraction/consumer.go:163
2Transcript + prompt (carrying the non-derivability rule) go to the extractorextraction/extractor.go:435
3Envelope parsed strictly — DisallowUnknownFieldsextraction/candidate.go:87
3aOne bad candidate rejects the whole envelope → retry → dead_letterextraction/consumer.go:237, ledger.go:140
4–5Provenance stamped, then deterministic classificationextraction/provenance.go:153, classify/constraint.go:173
6Redaction seam dispatches by classredact/detector.go:133
7Tier policy decides mutate-vs-flagredact/tier.go:252
8Gate escalates on any hit; otherwise the row activatesextraction/gate.go:150
9–10Human verdict, then deferred embed on approvalreview/reviewer.go:238
11The two vector recall paths filter on the embedder stampmcpsrv/pgbackend.go:224, assembly/memory_loader.go:277

2. Component catalog & status​

ComponentPathStatusNotes
Extraction enqueueextraction/enqueue.go:37🟢Gated on MEMORY_EXTRACTION_ENABLED
Extraction workerextraction/consumer.go:163🟢Redis consumer group
Provenance stampextraction/provenance.go:153🟢(model, endpoint class, prompt rev); AST-guarded
Deterministic classifierclassify/constraint.go:173🟢Two arms + context guards
Confidence gateextraction/candidate.go:155🟠Ships empty — cannot drop anything
Redaction seamredact/detector.go:133🟢REDACT_DETECTOR, code default regex
gliner NER hybridredact/nerdetector.go:146🟢ner in dev; sidecar + canary
National-ID checksumsredact/natid.go:174🟢Go-side, not the model
Tier policyredact/tier.go:252🟢Fail-closed zero value
Review gateextraction/gate.go:150🟢Escalate-only, tier-blind by design
Approval resolutionreview/approval_resolve.go:104🟢Voids paired approvals on terminal transitions
Sampled auditmemaudit/gate.go:115🟢10% (1000bp) — but see §8, no verdicts yet
Embed workermemory/embed_worker.go:217🟢
Stamp-scoped vector recall ×2see §1 step 11🟢The RAG path (retrieval/vector.go) got the same fix but queries context_embeddings, not memory
Stamp-exclusion counterrecall/stampcliff.go:71🟢Makes the mixed-corpus cliff visible
Near-dup collapsememory/neardup.go:167🟠OFF; threshold 0.10, ceiling 0.35
Prompt-Guard classifiercompose :945🟡ML_DETECTOR_USE_FAKE_CLASSIFIER=true — upstream repo is gated
StubExtractorextraction/extractor.go:59🔵Fallback when no API key resolves

3. The governance model — read this before changing anything​

Most of this subsystem's shape comes from rulings, not requirements. Code that looks arbitrary is usually a settled argument. Three documents govern:

ADR-0029 — PII detection stays inside the trust boundary​

Detection runs self-hosted, in-cluster; managed cloud DLP was rejected on BYOK grounds (a tenant who chose non-GCP would have had all session-derived content inspected by Google). Consequences that show up in code: the model artifact is baked and digest-pinned, Hub fetches are fatal, and the detector must work air-gapped.

It also established class-tiered mutation: quarantine already contains the risk, so mutating a quarantined row destroys the reviewer's evidence without protecting anything. The trigger was a real incident — a 13-digit Unix timestamp inside a code sample was matched as a phone number and irreversibly destroyed the artifact, because agent_memory has no pre-redaction original.

That trade rests on a premise (D3 (a)) which is a constraint on any future reader, not just a justification: a LOW-tier pending row holds un-redacted content at rest, and this is acceptable only because the reviewer is a human inside the owning org who already has access to the source session. Any surface that widens who can read a pending row — a new listing API, an export, a support tool, a cross-org view — breaks the premise the founder approved, not merely a convention.

ADR-0030 — model drops must be recorded​

A model classifier may remove a candidate before write only if the removal is recorded — countable, attributable to a model version, samplable. A drop with no record is forbidden. The harm being prevented is invisibility, not the model having an opinion: a drop filter's false positives are invisible by construction, because nothing counts what was never written.

Two consequences: attribution requires real provenance (hence §1 step 4 existing at all), and a model drop is admissible only on a candidate whose redaction class set is empty — redaction is an admissibility test, not a sanitiser.

⚠️ ADR-0030 is decided-only, not yet structural. Read D4 as a rule you must uphold, not one the code enforces for you. gate.go:91 still runs passesConfidence before RedactContent at :95 — so today the "no drop on a redaction-flagged candidate" invariant is held only because confidenceThresholds ships empty. The moment anyone populates that map without reordering those two calls, the ADR is violated in code while every test stays green. See §8 for the real prerequisites.

It also carries a residual the ADR names rather than closes (D5(d)): the rule binds only post-emission drops. An extractor prompt tightened to suppress candidates emits nothing to record, so defeating ADR-0030 in substance requires no code change at all — just a prompt edit. Anyone changing the extraction prompt is making a governance change, whether or not it looks like one.

The #2322 ruling — governance vs precision​

The extractor gets no vote on governance. high_stakes and importance are server-derived. confidence was carved out as a precision filter rather than a governance gate — a carve-out ADR-0030 later narrowed rather than overturned.


4. State machine & the review gate​

Every transition is CAS'd in updateStatus (review/reviewer.go:537); zero rows affected ⇒ ErrIllegalTransition.

From → ToOpNotes
pending → activeApproveEnqueues the deferred embed
active → activeApproveIdempotent, no re-embed
pending → rejectedReject
rejected → rejectedRejectIdempotent — this is the repair path
active|pending → tombstonedForget
active|pending → supersededCorrectChains via prev_memory_hash
active → supersedednear-dup collapseSets superseded_by; CAS on status='active'
rejected → tombstonedForgetILLEGAL — ErrIllegalTransition

Terminal: rejected, superseded, tombstoned. Only pending and active have outbound edges.

Every terminal transition — including the idempotent branches — also calls resolveOverride, voiding a still-pending paired approval with reason='operator_override_by_2507'. That idempotent branch is deliberate: re-running the operation is the repair for a stranded approval, which is why no one-off cleanup tool exists.

The gate itself (extraction/gate.go:150) is escalate-only and tier-blind by design: any redaction hit forces pending. It can raise scrutiny, never lower it.


5. Operate it​

#TaskToolEffect
1Review a memorycmd/tools/memory-review -op approve|reject|forget|correctTerminal transition + paired-approval resolution + audit row
2Re-run the classifier over a backlogcmd/tools/memory-status-reevaluate -dry-runRead-only decision diff; live pass needs founder approval
3Recover dead-lettered runscmd/tools/memory-extraction-replayStep 4 approval-repair is mandatory or recovered memories are invisible
4Repair pending rows with no approvalcmd/tools/memory-approval-repairFixes the #2323 stranding class
5Backfill embeddingscmd/tools/memory-backfillIdempotent; near-dup collapse forced OFF
6A/B a prompt changecmd/tools/memory-eval-abReal model calls — deliberately never in CI

Runbooks live under docs/runbooks/memory-*.md. Known defect: memory-review never exits after printing success (#2660) — a timeout wrapper returns 143 over already-applied work. Verify the DB, not the exit code.


6. Data model & API reference​

TableKey fieldsSource
agent_memorystatus, memory_type, content_hash, extractor_version, superseded_by, audited_at/audit_verdict/audit_sourcemig 005, 169, 196, 201
agent_memory_embeddingsembedding vector(1536), embed_model, HNSW cosinemig 170:50,75
governance_approvalsaction_type='memory_review', target=memory_id, reasonmig 173:32
agent_audit_logHash-chained per session; memory rows are session-lessmig 036, review/audit.go:65
context_embeddingsmodel_version NOT NULL, vector(1536)mig 004:47

Migrations in this arc: 169 (loop columns) · 170 (embeddings) · 171 (run ledger) · 173 (review gate) · 184 (replay) · 189 (importance default) · 191 (sampled audit) · 196 (superseded_by) · 197 (memory_reclassified) · 199 (orphaned-approval repair) · 200 (expiry backfill) · 201 (audited marker).

Surfaces: MCP recall_memory / remember (mcpsrv/tools.go:47,69) · PullContext with findLatest / findBySnapshot, both status='active'-scoped (service.go:517,550) · MLDetectorService.DetectPII (proto/upsquad/ml/v1/detector.proto:34).

status='active' is the same allow-list predicate on all seven recall surfaces — service.go:524,558 (PullContext's two arms), mcpsrv/pgbackend.go:58,218, assembly/memory_loader.go, runtime/session/warmstart.go:176, runtime/planmemory/recaller.go:127. All seven carry it, and an eighth must too — this was a real defect (ADR-0029 D5.2, where PullContext returned exactly the row that was supposed to be withheld).

Only two of the seven are vector searches, and only those carry the embed_model stamp filter (§7 trap 3): mcpsrv/pgbackend.go and assembly/memory_loader.go. The other five retrieve by id, recency or snapshot and have no vector to mis-compare — internal/runtime/ contains zero embed_model references, correctly. A new surface needs the stamp filter iff it does similarity search.

Metrics: 33 emitters under memmetrics, every one with a non-test caller, all label vocabularies bounded and enum-sourced. extractor_version labels must use Provenance.MetricLabel() (bounded), never Stamp().


7. Traps — the section that will save you a day​

1. Tiering is fail-closed by zero-value. Do not "helpfully" default a new class. TierFor is a bare map lookup; an absent key yields TierHigh (mutate). Enrolling a class as LOW is an explicit, reviewed act, and it adds a row to the BuildSummary pin. The tier table as shipped:

TierClassesBehaviour
LOWphone, ip, person, dob, address, detector_unavailableFlag + quarantine, content left verbatim
HIGH (by absence)secret, ssn, cc, email, natidMutate + quarantine

2. A quantised model detects nothing while reporting healthy. Measured during selection: the int8 gliner variant returned zero detections, health-green, and faster. That is why the shipped artifact is fp32 and why the health probe asserts a canary detection (a known-PII string must produce hits) rather than liveness — and why the compose healthcheck targets MLDetectorService.DetectPII, not the always-SERVING service key.

3. Never compare vectors across embed_model / model_version. Cosine geometry is not portable. All three recall paths filter on the live embedder's ModelVersion(). A mixed corpus without this mis-ranks with no error, no metric, nothing red. memory_recall_stamp_excluded_total exists so the exclusion is visible rather than looking like an empty corpus.

4. dead_letter is absorbing. claimRun drops terminal runs; only an explicit replay CAS re-arms one. A run that dead-letters stays dead until someone runs the replay tool — and the replay runbook's step 4 (approval repair) is mandatory, or the recovered memories exist but are invisible to reviewers.

5. A bad candidate rejects the whole envelope. ParseEnvelope uses DisallowUnknownFields and fails the entire run on one malformed candidate. This is why the non-derivability rule rides the existing Rationale field rather than a new required field: a new required field would convert one model formatting lapse into total memory loss for that session.

6. Two retention mechanisms, and only one of them works on loop-written memories. The org-level RTD sweeper (4h tick, agent_memory at 180d default / 30d floor) does reach these rows — they are not immortal. But the per-row expires_at / retention_days columns are written only by PushContext; the extraction, remember and Correct INSERTs omit them, so every loop-written memory has expires_at NULL and is invisible to idx_memory_expires_at. Per-memory_type decay (AML-F16) is genuinely not built.

7. Comments that disagree with the code (correct as of d3d8886e, fix them when you touch them):

WhereThe lie
redact/detector.go:12-20Says the seam has no NER implementation and ner falls back to regex. False since #2155 — only the dlp arm still matches
extraction/gate.go:1-6Reads as three live gates; the confidence gate cannot drop (empty map)
extraction/extractor.go:4-6Calls StubExtractor the default wiring; the composition root prefers GeminiExtractor whenever a key resolves

8. Code default ≠ deployed default. MEMORY_EXTRACTION_ENABLED is OFF in code, true in compose. Reading only one will mislead you.

EMBEDDING_PROVIDER used to be the worst instance of this — openai in code, vertex in compose, so which third party received a tenant's documents depended on which layer you read (ADR-0029). Both defaults are deleted by LLD #2714 T4 (PRD #996 FR-9 step 1, #2546): unset is now a refusal and a non-zero exit with a message naming the choices, and docker-compose.dev.yml states vertex outright rather than as a ${...:-vertex} fallback. There is no longer a layer to read it off; the only answer is the one in the manifest and the one on the boot line.


8. Gaps, pending, parked — the honest state​

ItemRefState
Prod extraction enablementADR-0029 D4.4🟠 OFF until C1–C3 below
↳ C1 live dogfood run under the shipped prompt—pending
↳ C2 human-rated precision sample—pending
↳ C3 extraction spend metered#2669🟠 prod-enablement — no llm_usage_events, spend unbillable and uncappable
Sampled audit produces 0 decided verdicts in ~4 months#2654🟠 "samplable but not sampled" — the state ADR-0030 D2 rejects
Confidence gate inert (empty thresholds)candidate.go:155🟠 Do not populate this map casually. It is what currently holds ADR-0030 D4 — see below
Near-dup collapseneardup.go:167🟠 OFF; live active corpus has no true near-duplicates at any safe threshold
Prompt-Guard classifier is a fakecompose :945🟡 Upstream repo gated
Detector image 9.21 GB (CUDA wheels, CPU-only service)#2666🟠 Defeats the footprint rationale of the model choice
Per-memory_type decay (AML-F16)—🟠 Not built; blocks the PRD §10 decay/retention tier claims
Tier consolidation (AML-F18)—🟠 Phase 3, not due
context_shares empty OrgIDColumnscopes.go:658🟠 Opts out of the generic deleter
memory-review never exits#2660🟠 Zero test files in that package

What enabling the confidence gate actually costs​

The single most likely way to break ADR-0030 is to "just add a threshold". Populating confidenceThresholds turns the model into a live drop gate, which the ADR permits only once all of D1–D4 hold. The full prerequisite set, not the label count alone:

#PrerequisiteState
D1A drop-record store that retains the dropped content — a counter is not "samplable", and memory_candidates_dropped_total already exists, so counter-only would rename the status quo as compliancenot built
D2A mandated non-zero sample rate over that store — "samplable but not sampled" is explicitly rejectednot built; and #2654 shows the existing audit stream at 0 verdicts
D3Real provenance so drops are attributable✅ shipped (#2584)
D4The drop evaluated after redaction, so it can never remove a flagged candidate — gate.go:91 still runs it before :95not done
—A classregistry scope for the new retained-content store (class, retention, OrgIDColumn, working Deleter)not built
—P0.1.3 DPA widening — retained never-promoted content is a new purpose and category, and DSAR/erasure must reach itnot built

The "≥11 verdicts rendered against the current predicate" figure is the floor-promotion condition for scoring a stamp — a much later and smaller question than switching the gate on.


9. Key files, flags & environment​

Flags (code default → dev value)

VarCode defaultbeta-dev
MEMORY_EXTRACTION_ENABLEDOFFtrue
REDACT_DETECTORregexner
ML_DETECTOR_ADDRESSk8s DNSml-detector:50060
EMBEDDING_PROVIDERnone — refuses to boot (#2714 T4)vertex (stated in compose, not defaulted)
EMBEDDING_DIM15361536
MEMORY_NEAR_DUP_COLLAPSE_ENABLEDOFFfalse
MEMORY_NEAR_DUP_MAX_DISTANCE0.10 (ceiling 0.35, refuses rather than clamps)unset
MEMORY_EXTRACTION_TIMEOUT30s120s

EMBEDDING_DIM configures only the embedder client. Both vector(1536) literals are unreachable from it — changing it produces a runtime insert error, not a schema change.

Source map — extraction internal/context/memory/extraction/ · classification .../classify/ · redaction + tiering .../redact/ · review .../review/ · recall .../mcpsrv/, internal/context/assembly/, internal/context/retrieval/ · detector sidecar services/ml-detector/ · CLIs cmd/tools/memory-*.

CI lanes — go-test.yml (unconditional, carries the eval floors) · db.yml → Migrate Up + Down (six -tags integration memory steps, per-issue anti-skip needles) · ml-detector-tests.yml (pytest + the real-model corpus floor) · agent-worker-tests.yml. The last two exist because those suites had never run in any workflow.

Related: ADR-0029 · ADR-0030 · runbooks docs/runbooks/memory-*.md · PRD #2037 · tracker #2063.


Doc reflects upsquad-core@d3d8886e (2026-08-29). Invalidated by: prod enablement flipping, #2654 producing verdicts (the confidence gate becomes promotable), #2666 changing the artifact, or any new recall surface — which must carry both the status='active' and embed_model predicates.