Memory extraction — dead-letter replay
Issue #2308 · milestone #27 (Agent Memory Loop, PRD #2037) · tool:
cmd/tools/memory-extraction-replay· ledger:memory_extraction_runs
Replay is two tools, not one.
memory-extraction-replayrecovers the extraction; it does not open the governance approvals that quarantined memories need.cmd/tools/memory-approval-repairis step 4 below and it is mandatory — skipping it leaves recovered memories persisted, embedded and permanently invisible to the reviewer (#2383). Do not stop at step 3.
When you need this
An extraction run that fails 3 times lands in memory_extraction_runs.state = 'dead_letter'. Before #2308 that was permanent: the consumer acks the job so
Redis never redelivers it, and claimRun drops any later delivery for a terminal
run. The session's memories were simply gone.
On beta-dev 2026-07-28, four runs dead-lettered inside one ~2-minute window from
pure local-model contention — runs either side of them completed in 13.1s and
10.6s against a 30s cap. Day tally: 11 succeeded, 4 dead_letter. Nothing was
structurally wrong; a brief load spike cost four sessions their memories.
The source transcript is still in session_events. Only the attempt was lost.
This runbook is the path back.
Related: #2298 / PR #2306 makes the extraction timeout tunable, which reduces how often this triggers. It does not address what to do when it does.
Detect
-- dead letters by day and failure mode
SELECT date_trunc('day', updated_at) AS day,
count(*) FILTER (WHERE state = 'dead_letter') AS dead,
count(*) FILTER (WHERE state = 'succeeded') AS ok
FROM memory_extraction_runs
GROUP BY 1 ORDER BY 1 DESC LIMIT 14;
-- what actually killed them
SELECT session_id, org_id, attempts, replays, last_replay_at, left(last_error, 160)
FROM memory_extraction_runs
WHERE state = 'dead_letter'
ORDER BY updated_at DESC LIMIT 50;
Metrics (Context Engine, OTel → Prometheus):
| Metric | Meaning |
|---|---|
memory_extraction_dlq_total{org_id} | runs dead-lettered after the retry cap |
memory_extraction_replay_total{org_id,outcome} | deliberate replays — recovered / failed / skipped / no_transcript |
increase(memory_extraction_dlq_total[1h]) is the alert signal. The permanent
loss is dlq_total − replay_total{outcome="recovered"}.
Procedure
Five steps. Steps 4 and 5 are not optional — a replay that stops at step 3 is a half-recovery (see "Replay does not open approvals" below).
1. Environment
export DATABASE_URL=postgres://... # a role that reads across orgs (owner/BYPASSRLS)
export REDIS_ADDR=redis:6379 # for the embed-on-write enqueue
export REDIS_PASSWORD=... # only if the instance requires auth
export MEMORY_EXTRACTION_API_KEY=... # or COMPACTION_API_KEY
export MEMORY_EXTRACTION_ENDPOINT=... # optional; else COMPACTION_ENDPOINT, else default
export AGENT_ORCHESTRATOR_GRPC_ADDR=agent-orchestrator:50051 # required by step 4
REDIS_ADDR + REDIS_PASSWORD, never REDIS_URL. memory-extraction-replay
and memory-backfill both read the _ADDR/_PASSWORD pair
(cmd/tools/memory-extraction-replay/main.go, openRedis). A REDIS_URL you
exported is ignored in silence and the tool falls back to localhost:6379, which
on a devbox is a different Redis than the stack's.
AGENT_ORCHESTRATOR_GRPC_ADDR is required by step 4, not by the replay
itself. memory-approval-repair fails fast without it — "the ApprovalService is
constructed exactly once, in agent-orchestrator". Export it up front so you do
not discover it after the replay has already run.
On the devbox, run both tools inside the compose network. litellm publishes
no host port, so a host-run replay cannot reach the extractor endpoint at all,
and agent-orchestrator is only agent-orchestrator:50051 from inside. Attach to
the network rather than translating every address:
docker run --rm --network upsquad-dev_upsquad-dev \
-v "$PWD":/src -w /src \
-e DATABASE_URL -e REDIS_ADDR -e REDIS_PASSWORD \
-e MEMORY_EXTRACTION_API_KEY -e AGENT_ORCHESTRATOR_GRPC_ADDR \
golang:1.25 go run ./cmd/tools/memory-extraction-replay -dry-run
(From the host instead, agent-orchestrator is published on 127.0.0.1:50052
and redis on 127.0.0.1:6379 — but litellm is not published at all, so the
host route works for step 4 and never for step 3.)
2. Dry run
Always dry-run first. It costs nothing, calls no model, and writes nothing.
go run ./cmd/tools/memory-extraction-replay -dry-run
3. Replay
# Replay transient failures from the last 24h, at most 20 of them
go run ./cmd/tools/memory-extraction-replay -max-age 24h -limit 20
# One specific session, whatever the failure reason
go run ./cmd/tools/memory-extraction-replay -session <uuid> -include-nonretryable
| Flag | Default | Notes |
|---|---|---|
-dry-run | off | list only; no model call, no writes. Exempt from the -yes requirement |
-max-age | 168h | horizon on enqueued_at; 0 disables. Replaying a months-old transcript against today's prompt produces memories nobody reviewed in context |
-limit | 25 | caps the blast radius — each replay is one model call. 0 means no limit and additionally requires -yes |
-yes | off | confirms an unbounded run (-limit 0): every eligible dead letter in every org |
-max-replays | 3 | poison cap; a run replayed 3× is not a blip |
-batch | 50 | dead-letter rows fetched per keyset page |
-org / -session | all | narrow the sweep |
-include-nonretryable | off | also replay failures not classified transient (parse errors, malformed envelopes) |
-force | off | re-extract even when agent_memory already holds live rows for the session. Risk: duplicates unless the extractor re-emits byte-identical content |
Outcomes in the summary line:
| Outcome | Meaning |
|---|---|
recovered | re-extracted and the run reached succeeded; memories are in agent_memory |
failed | failed again; row is back in dead_letter with the new error, replays incremented |
skipped | not dead-lettered any more, or at the replay cap. Nothing changed |
no_transcript | session_events is empty — nothing to recover. Row left dead-lettered on purpose, so the loss stays visible |
no_agent | the session row resolves to no agent — nothing to persist against. Left dead-lettered |
already_written | agent_memory already holds live rows for the session (the run persisted, then failed before markSucceeded). Left dead-lettered; re-extracting would duplicate. -force overrides |
4. Repair the approvals — MANDATORY
Run this after every replay that recovered anything. A recovered count
greater than zero is not a completed recovery.
# 4a. Dry run. Needs only DATABASE_URL; contacts no governance surface.
go run ./cmd/tools/memory-approval-repair -dry-run
# 4b. Repair. Narrow the first live run to the ids the dry run printed.
go run ./cmd/tools/memory-approval-repair -org <org-uuid> [-memory <uuid,uuid>]
| Flag | Default | Notes |
|---|---|---|
-dry-run | off | enumerate candidates only; opens no approval, needs no governance connection |
-org | all | narrows the predicate and sets app.org_id on the read tx |
-agent | all | narrows the predicate and sets app.agent_id on the read tx |
-memory | all | comma-separated memory ids. A filter, never an override |
-batch | 200 | keyset page size |
-limit | 0 | max rows to repair; 0 = no limit |
-deadline | 0 | review window from now. 0 inherits the review-gate default (7d). Must be positive — never backdate it, an expired approval is treated as a denial and rejects the memory you are recovering |
The summary line is candidates=N repaired=N failed=0. failed > 0 exits
non-zero; the failed rows stay in the predicate, so re-run to retry.
Full detail, RLS notes and failure modes:
memory-approval-repair.md.
5. Verify zero stranded rows
This is the check that would have caught the 2026-07-31 incident. It must return
0.
-- Stranded: quarantined memories with NO memory-review approval row of any status.
-- This is exactly the repair tool's predicate.
SELECT count(*) AS stranded
FROM agent_memory m
LEFT JOIN governance_approvals g
ON g.action_type = 'memory_review'
AND (g.target = m.id::text OR g.metadata->>'memory_id' = m.id::text)
WHERE m.status = 'pending'
AND g.id IS NULL;
To see which rows, and who owns them:
SELECT m.id, m.memory_type, m.created_at, a.name AS agent
FROM agent_memory m
LEFT JOIN governance_approvals g
ON g.action_type = 'memory_review'
AND (g.target = m.id::text OR g.metadata->>'memory_id' = m.id::text)
LEFT JOIN agents a ON a.id = m.agent_id
WHERE m.status = 'pending'
AND g.id IS NULL
ORDER BY m.created_at;
Cluster the results by created_at — a tight window almost always maps to one
incident window.
Schema traps in these queries. Each of these cost real time on 2026-07-31:
| Trap | Reality |
|---|---|
| Join key | governance_approvals.target = agent_memory.id::text, with action_type = 'memory_review'. The requester writes the id to both target and metadata->>'memory_id'; match either, as the tool does |
governance_approvals.agent_id | TEXT, while agents.id is UUID. Joining the two needs an explicit cast: a.id::text = ga.agent_id. (Joining via agent_memory.agent_id, as above, needs no cast — that column is UUID) |
governance_approvals has no created_at | The timestamp column is requested_at. resolved_at / expires_at are the others |
memory_extraction_runs has no status | The state column is state (enqueued/running/succeeded/failed/dead_letter) |
memory_extraction_runs has no created_at | Use updated_at, or enqueued_at for when the run was first queued |
Replay does not open approvals — and that is deliberate
memory-extraction-replay drives the identical online extraction path
(transcript → extract → gate → redact → persist → markSucceeded → embed
enqueue). The embed seam is wired: persisted candidates enqueue onto the global
memory:embed stream exactly as the online consumer does.
The review-requester side is deliberately left unwired
(cmd/tools/memory-extraction-replay/main.go):
The review-requester side is left unwired (pre-PR-6 behaviour): a quarantined replayed memory is persisted and inert, and opening governance approvals from a CLI tool would fan out notifications from an operator's terminal.
That is a design choice, not a bug: an operator running a bulk recovery at 02:00
should not page every reviewer in the org from their shell. Do not "fix" it by
wiring the requester into the replay tool. The completion step is step 4 — a
separate, explicit, idempotent tool that opens the approvals through the same
review.Requester the online consumer uses, so template, required clearance and
expiry match by construction.
The consequence of skipping step 4 is total: a quarantined memory
(status='pending') with no governance_approvals row is in an absorbing state.
The approvals inbox lists approval rows, so it never renders. RecordDecision
keys on an approval id that does not exist. The memory_extraction_runs ledger
makes re-extraction a no-op. It can never be approved, so it can never activate.
This has happened. On 2026-07-31 a replay of 28 dead letters reported
recovered=20 failed=8 and looked like a success. It left 24 memories
pending with no approval row (21 from that replay, 3 pre-existing) — invisible
to the reviewer. memory-approval-repair returned candidates=24 repaired=24 failed=0 and the stranded count went to 0. Nothing was lost, but only because
someone ran the join by hand.
Why it is safe to run
- It cannot duplicate a good run. The only exit from
dead_letteris a CAS guarded onstate = 'dead_letter', so asucceededrun is structurally unreachable. The online consumer's claim path is unchanged. - It cannot double-run. The CAS is one
UPDATE ... WHERE; concurrent invocations produce exactly one winner. - It cannot loop forever.
replays < -max-replays. - It is resumable. Work is found by DB predicate and keyset-paged, so interrupting and re-running resumes; recovered runs drop out of the predicate.
- It refuses to run blind. Without a real extractor configured it exits
non-zero rather than stamping every run
succeededwith zero candidates.
Gotchas
- Run it as an admin DB role. Like
cmd/tools/memory-backfillit reads across orgs. Every write re-derivesapp.org_idfrom the row, so writes stay scope-correct under FORCE RLS regardless of the read role. - A replay is a fresh model call, not a stored intermediate — the extractor
output of a failed run is never persisted. Candidates may differ from what the
original run would have produced; the persist path dedups on redacted
content_hash, so overlap converges to a no-op. - Fix the cause first. If dead letters are still arriving (a wedged upstream,
a too-tight
MEMORY_EXTRACTION_TIMEOUT), replaying just re-burns quota. Replay once the spike is over. - A clean
recovered=Nsummary is not "done". The replay tool reports on the extraction ledger only; it has no visibility into whether the memories it wrote are reachable by a reviewer. Step 4 + step 5 are the completion check, and the only surface that will tell you.
Related
memory-approval-repair.md— the step-4 tool in full: RLS notes, idempotency and dedup properties, why a repaired approval is never backdated.memory-summary-backfill.md— the sibling DB-predicate-driven repair for missing summaries.