Skip to main content

Memory extraction — dead-letter replay

Issue #2308 · milestone #27 (Agent Memory Loop, PRD #2037) · tool: cmd/tools/memory-extraction-replay · ledger: memory_extraction_runs

Replay is two tools, not one. memory-extraction-replay recovers the extraction; it does not open the governance approvals that quarantined memories need. cmd/tools/memory-approval-repair is step 4 below and it is mandatory — skipping it leaves recovered memories persisted, embedded and permanently invisible to the reviewer (#2383). Do not stop at step 3.

When you need this​

An extraction run that fails 3 times lands in memory_extraction_runs.state = 'dead_letter'. Before #2308 that was permanent: the consumer acks the job so Redis never redelivers it, and claimRun drops any later delivery for a terminal run. The session's memories were simply gone.

On beta-dev 2026-07-28, four runs dead-lettered inside one ~2-minute window from pure local-model contention — runs either side of them completed in 13.1s and 10.6s against a 30s cap. Day tally: 11 succeeded, 4 dead_letter. Nothing was structurally wrong; a brief load spike cost four sessions their memories.

The source transcript is still in session_events. Only the attempt was lost. This runbook is the path back.

Related: #2298 / PR #2306 makes the extraction timeout tunable, which reduces how often this triggers. It does not address what to do when it does.

Detect​

-- dead letters by day and failure mode
SELECT date_trunc('day', updated_at) AS day,
count(*) FILTER (WHERE state = 'dead_letter') AS dead,
count(*) FILTER (WHERE state = 'succeeded') AS ok
FROM memory_extraction_runs
GROUP BY 1 ORDER BY 1 DESC LIMIT 14;

-- what actually killed them
SELECT session_id, org_id, attempts, replays, last_replay_at, left(last_error, 160)
FROM memory_extraction_runs
WHERE state = 'dead_letter'
ORDER BY updated_at DESC LIMIT 50;

Metrics (Context Engine, OTel → Prometheus):

MetricMeaning
memory_extraction_dlq_total{org_id}runs dead-lettered after the retry cap
memory_extraction_replay_total{org_id,outcome}deliberate replays — recovered / failed / skipped / no_transcript

increase(memory_extraction_dlq_total[1h]) is the alert signal. The permanent loss is dlq_total − replay_total{outcome="recovered"}.

Procedure​

Five steps. Steps 4 and 5 are not optional — a replay that stops at step 3 is a half-recovery (see "Replay does not open approvals" below).

1. Environment​

export DATABASE_URL=postgres://... # a role that reads across orgs (owner/BYPASSRLS)
export REDIS_ADDR=redis:6379 # for the embed-on-write enqueue
export REDIS_PASSWORD=... # only if the instance requires auth
export MEMORY_EXTRACTION_API_KEY=... # or COMPACTION_API_KEY
export MEMORY_EXTRACTION_ENDPOINT=... # optional; else COMPACTION_ENDPOINT, else default
export AGENT_ORCHESTRATOR_GRPC_ADDR=agent-orchestrator:50051 # required by step 4

REDIS_ADDR + REDIS_PASSWORD, never REDIS_URL. memory-extraction-replay and memory-backfill both read the _ADDR/_PASSWORD pair (cmd/tools/memory-extraction-replay/main.go, openRedis). A REDIS_URL you exported is ignored in silence and the tool falls back to localhost:6379, which on a devbox is a different Redis than the stack's.

AGENT_ORCHESTRATOR_GRPC_ADDR is required by step 4, not by the replay itself. memory-approval-repair fails fast without it — "the ApprovalService is constructed exactly once, in agent-orchestrator". Export it up front so you do not discover it after the replay has already run.

On the devbox, run both tools inside the compose network. litellm publishes no host port, so a host-run replay cannot reach the extractor endpoint at all, and agent-orchestrator is only agent-orchestrator:50051 from inside. Attach to the network rather than translating every address:

docker run --rm --network upsquad-dev_upsquad-dev \
-v "$PWD":/src -w /src \
-e DATABASE_URL -e REDIS_ADDR -e REDIS_PASSWORD \
-e MEMORY_EXTRACTION_API_KEY -e AGENT_ORCHESTRATOR_GRPC_ADDR \
golang:1.25 go run ./cmd/tools/memory-extraction-replay -dry-run

(From the host instead, agent-orchestrator is published on 127.0.0.1:50052 and redis on 127.0.0.1:6379 — but litellm is not published at all, so the host route works for step 4 and never for step 3.)

2. Dry run​

Always dry-run first. It costs nothing, calls no model, and writes nothing.

go run ./cmd/tools/memory-extraction-replay -dry-run

3. Replay​

# Replay transient failures from the last 24h, at most 20 of them
go run ./cmd/tools/memory-extraction-replay -max-age 24h -limit 20

# One specific session, whatever the failure reason
go run ./cmd/tools/memory-extraction-replay -session <uuid> -include-nonretryable
FlagDefaultNotes
-dry-runofflist only; no model call, no writes. Exempt from the -yes requirement
-max-age168hhorizon on enqueued_at; 0 disables. Replaying a months-old transcript against today's prompt produces memories nobody reviewed in context
-limit25caps the blast radius — each replay is one model call. 0 means no limit and additionally requires -yes
-yesoffconfirms an unbounded run (-limit 0): every eligible dead letter in every org
-max-replays3poison cap; a run replayed 3× is not a blip
-batch50dead-letter rows fetched per keyset page
-org / -sessionallnarrow the sweep
-include-nonretryableoffalso replay failures not classified transient (parse errors, malformed envelopes)
-forceoffre-extract even when agent_memory already holds live rows for the session. Risk: duplicates unless the extractor re-emits byte-identical content

Outcomes in the summary line:

OutcomeMeaning
recoveredre-extracted and the run reached succeeded; memories are in agent_memory
failedfailed again; row is back in dead_letter with the new error, replays incremented
skippednot dead-lettered any more, or at the replay cap. Nothing changed
no_transcriptsession_events is empty — nothing to recover. Row left dead-lettered on purpose, so the loss stays visible
no_agentthe session row resolves to no agent — nothing to persist against. Left dead-lettered
already_writtenagent_memory already holds live rows for the session (the run persisted, then failed before markSucceeded). Left dead-lettered; re-extracting would duplicate. -force overrides

4. Repair the approvals — MANDATORY​

Run this after every replay that recovered anything. A recovered count greater than zero is not a completed recovery.

# 4a. Dry run. Needs only DATABASE_URL; contacts no governance surface.
go run ./cmd/tools/memory-approval-repair -dry-run

# 4b. Repair. Narrow the first live run to the ids the dry run printed.
go run ./cmd/tools/memory-approval-repair -org <org-uuid> [-memory <uuid,uuid>]
FlagDefaultNotes
-dry-runoffenumerate candidates only; opens no approval, needs no governance connection
-orgallnarrows the predicate and sets app.org_id on the read tx
-agentallnarrows the predicate and sets app.agent_id on the read tx
-memoryallcomma-separated memory ids. A filter, never an override
-batch200keyset page size
-limit0max rows to repair; 0 = no limit
-deadline0review window from now. 0 inherits the review-gate default (7d). Must be positive — never backdate it, an expired approval is treated as a denial and rejects the memory you are recovering

The summary line is candidates=N repaired=N failed=0. failed > 0 exits non-zero; the failed rows stay in the predicate, so re-run to retry.

Full detail, RLS notes and failure modes: memory-approval-repair.md.

5. Verify zero stranded rows​

This is the check that would have caught the 2026-07-31 incident. It must return 0.

-- Stranded: quarantined memories with NO memory-review approval row of any status.
-- This is exactly the repair tool's predicate.
SELECT count(*) AS stranded
FROM agent_memory m
LEFT JOIN governance_approvals g
ON g.action_type = 'memory_review'
AND (g.target = m.id::text OR g.metadata->>'memory_id' = m.id::text)
WHERE m.status = 'pending'
AND g.id IS NULL;

To see which rows, and who owns them:

SELECT m.id, m.memory_type, m.created_at, a.name AS agent
FROM agent_memory m
LEFT JOIN governance_approvals g
ON g.action_type = 'memory_review'
AND (g.target = m.id::text OR g.metadata->>'memory_id' = m.id::text)
LEFT JOIN agents a ON a.id = m.agent_id
WHERE m.status = 'pending'
AND g.id IS NULL
ORDER BY m.created_at;

Cluster the results by created_at — a tight window almost always maps to one incident window.

Schema traps in these queries. Each of these cost real time on 2026-07-31:

TrapReality
Join keygovernance_approvals.target = agent_memory.id::text, with action_type = 'memory_review'. The requester writes the id to both target and metadata->>'memory_id'; match either, as the tool does
governance_approvals.agent_idTEXT, while agents.id is UUID. Joining the two needs an explicit cast: a.id::text = ga.agent_id. (Joining via agent_memory.agent_id, as above, needs no cast — that column is UUID)
governance_approvals has no created_atThe timestamp column is requested_at. resolved_at / expires_at are the others
memory_extraction_runs has no statusThe state column is state (enqueued/running/succeeded/failed/dead_letter)
memory_extraction_runs has no created_atUse updated_at, or enqueued_at for when the run was first queued

Replay does not open approvals — and that is deliberate​

memory-extraction-replay drives the identical online extraction path (transcript → extract → gate → redact → persist → markSucceeded → embed enqueue). The embed seam is wired: persisted candidates enqueue onto the global memory:embed stream exactly as the online consumer does.

The review-requester side is deliberately left unwired (cmd/tools/memory-extraction-replay/main.go):

The review-requester side is left unwired (pre-PR-6 behaviour): a quarantined replayed memory is persisted and inert, and opening governance approvals from a CLI tool would fan out notifications from an operator's terminal.

That is a design choice, not a bug: an operator running a bulk recovery at 02:00 should not page every reviewer in the org from their shell. Do not "fix" it by wiring the requester into the replay tool. The completion step is step 4 — a separate, explicit, idempotent tool that opens the approvals through the same review.Requester the online consumer uses, so template, required clearance and expiry match by construction.

The consequence of skipping step 4 is total: a quarantined memory (status='pending') with no governance_approvals row is in an absorbing state. The approvals inbox lists approval rows, so it never renders. RecordDecision keys on an approval id that does not exist. The memory_extraction_runs ledger makes re-extraction a no-op. It can never be approved, so it can never activate.

This has happened. On 2026-07-31 a replay of 28 dead letters reported recovered=20 failed=8 and looked like a success. It left 24 memories pending with no approval row (21 from that replay, 3 pre-existing) — invisible to the reviewer. memory-approval-repair returned candidates=24 repaired=24 failed=0 and the stranded count went to 0. Nothing was lost, but only because someone ran the join by hand.

Why it is safe to run​

  • It cannot duplicate a good run. The only exit from dead_letter is a CAS guarded on state = 'dead_letter', so a succeeded run is structurally unreachable. The online consumer's claim path is unchanged.
  • It cannot double-run. The CAS is one UPDATE ... WHERE; concurrent invocations produce exactly one winner.
  • It cannot loop forever. replays < -max-replays.
  • It is resumable. Work is found by DB predicate and keyset-paged, so interrupting and re-running resumes; recovered runs drop out of the predicate.
  • It refuses to run blind. Without a real extractor configured it exits non-zero rather than stamping every run succeeded with zero candidates.

Gotchas​

  • Run it as an admin DB role. Like cmd/tools/memory-backfill it reads across orgs. Every write re-derives app.org_id from the row, so writes stay scope-correct under FORCE RLS regardless of the read role.
  • A replay is a fresh model call, not a stored intermediate — the extractor output of a failed run is never persisted. Candidates may differ from what the original run would have produced; the persist path dedups on redacted content_hash, so overlap converges to a no-op.
  • Fix the cause first. If dead letters are still arriving (a wedged upstream, a too-tight MEMORY_EXTRACTION_TIMEOUT), replaying just re-burns quota. Replay once the spike is over.
  • A clean recovered=N summary is not "done". The replay tool reports on the extraction ledger only; it has no visibility into whether the memories it wrote are reachable by a reviewer. Step 4 + step 5 are the completion check, and the only surface that will tell you.