Runbook — changing the embedding dimensionality of an install
Owner: backend-sme · Applies to: every install past migration 210 Refs: LLD #2714 §D1 · PRD #996 FR-6/FR-7 · ADR-0031 · issue 2727
The one thing to read before anything else
Schema rollback is real and takes seconds. Data rollback does not exist and never
did. Reversing the schema means re-embedding again. There is no cross-arity
projection that preserves meaning, so nothing anywhere in this system can give
you your old vectors back once you have confirmed the change. The sources —
rag_chunks and agent_memory — are never touched by any step below, which is
what makes a re-embed a recovery rather than a loss.
Founder decision FD-4 (PRD #996) says a re-embed is always acceptable. This runbook is written on that basis.
What decides the arity
One number per install (FD-3). Three things must agree, and the system now refuses to run when they do not:
| Fact | Where it lives | What disagreement causes |
|---|---|---|
| What the embedder produces | probed from the live endpoint at boot (Dimensions()) | — |
| What the schema can hold | context_embeddings.embedding, agent_memory_embeddings.embedding | FR-7 refuses to boot, naming both numbers |
| What the migration targets | the upsquad.embedding_dim connection GUC | the migration no-ops or refuses; never silently half-applies |
EMBEDDING_DIM is a claim about a server. The embedder probes the server
and the probe wins. If your config says 768 and the endpoint serves 1536, the
process fails at construction with ErrServedDimensionMismatch — before FR-7
is even reached.
Procedure — moving an install to a new arity
Order matters. ALTER … TYPE takes ACCESS EXCLUSIVE, and a row of the old arity
present makes it fail rather than corrupt, so the migration cannot be
interleaved with a re-embed.
1. Stop the writers
context-engine, upsquad-memory-mcp, and any running memory-backfill. They
will refuse to start again until step 3, by design.
2. Re-dimension the schema
migrate -path internal/context/store/migrations \
-database 'postgres://…/upsquad?sslmode=disable&options=-c%20upsquad.embedding_dim%3D768%20-c%20upsquad.embedding_redim_confirm%3Dyes' \
up
%20 is a space and %3D is = — the whole options= value rides inside a
URL query parameter.
On the dev compose stack, set both in .env instead and run make dev-up once:
EMBEDDING_DIM=768
EMBEDDING_REDIM_CONFIRM=yes
Then remove EMBEDDING_REDIM_CONFIRM again. Leaving it set means the next
person who edits EMBEDDING_DIM clears two corpora by starting a stack.
What you will see:
| You passed | What happens |
|---|---|
| neither GUC | no-op, rc=0, every row retained. This is the normal path for every migration harness in the repo. |
embedding_dim == the live arity | no-op, every row retained — even with the confirmation set. |
embedding_dim != live, no confirmation | refusal, naming both arities and the row count. Nothing changes. |
| both | both tables truncated and re-typed in one transaction. |
3. Point the services at the new embedder
Set EMBEDDING_PROVIDER / EMBEDDING_BASE_URL / EMBEDDING_MODEL /
EMBEDDING_DIM and start them. The FR-7 check runs at boot; if the schema and
the embedder disagree you get a message naming both numbers and a non-zero exit,
not a pgvector error hours later.
4. Flush the embedding cache (belt and braces)
redis-cli --scan --pattern 'emb:*' | xargs -r redis-cli del
Postgres TRUNCATE does not reach Redis, and emb:{content_hash}:{model_version}
has a 30-day TTL. This step is not load-bearing and you should know why:
since LLD #2714 T1/T2 the model identity carries the arity, so a re-dimension
necessarily changes model_version and a post-migration read cannot hit a
pre-migration entry. The flush is for disk, not for correctness — which is the
right way round, because a runbook step is a step someone skips.
5. Re-embed
Both corpora are empty until you do. Recall does not error in the meantime: retrieval falls back to BM25 + recency and memory recall returns empty.
memory-backfill --batch 200 # agent memory
rag-reembed --batch 200 # rag_chunks / context_embeddings
Both take --dry-run (count only, writes nothing) and --limit N. Run the dry
run first: the number it prints is the real backlog under the identity this
process resolved, not an estimate — both tools construct their embedder BEFORE
counting, precisely so the count is a claim about a known identity.
Both are resumable by re-running them. There is no resume file and no bookmark: a row that has been re-embedded no longer satisfies the candidate predicate, so an interrupted run restarted from the beginning walks only what is left. Interrupt them freely.
Watch the final line of each. converged=true means every candidate the run
identified was written; anything else means re-run. A run that ends
converged=false with a non-zero failed has skipped individual bad rows and
kept going — it will attempt each of them exactly once more per run, so re-run
and then investigate whatever still fails.
rag-reembed does not delete rows written under a previous
model_version. Recall excludes them already (retrieval/vector.go filters the
ANN on the live stamp), and keeping them is what makes a rollback to the previous
embedder free rather than a second full re-embed. Migration 211's
(chunk_id, model_version) uniqueness bounds that to one row per chunk per model
you have actually used. If you are reclaiming disk after a switch you are certain
of, that DELETE is a deliberate operator action, not something the tool does
unattended.
Check convergence from the outside too. GetSourceStatus reports
pending_chunk_count per knowledge source against the live embedder identity —
that number going to zero is the operator-visible form of the same fact. Before
LLD #2714 T7 it counted any embedding row and therefore read 0 whatever had
happened, which is why this step used to look complete the moment it started.
Recovering from an embedding paused degradation (MG-CTX.2)
rag-reembed is the ONLY recovery, and nothing re-drives those documents for
you. A re-dimension is not the only way rag_chunks rows come to exist with no
vector. Since LLD #3138
task 5D3 (#3163), an
embedding capability that cannot serve — credential absent or expired, endpoint
gone, breaker open — puts the capability in the declared embedding paused state
and the ingest continues: the chunk rows land, no context_embeddings row
does, and recall falls back to BM25 + recency. That is deliberate (the
alternative rolled back the whole ingest, taking the source's bookkeeping and
audit row with it), and it leaves exactly the condition step 5 above repairs.
What it does not leave is a queue. Configuring a live binding fixes the NEXT ingest and repairs nothing already landed:
- knowledge ingest is hash-idempotent — an unchanged body hash is a no-op, so a later sync never re-drives a document ingested while paused;
- the conversation embed-worker acks a degraded run, because the degradation is a success by design.
So after fixing the credential or the binding:
rag-reembed --dry-run # the backlog, under the identity this process resolved
rag-reembed --batch 200 # then the repair; re-runnable, no bookmark
The number to watch is the same one as above — each source's
pending_chunk_count — and the fleet-wide view is
context_capability_degraded_work_total{capability="embedding",unit="chunk"},
which counts every chunk left unembedded rather than one per incident.
Rolling back
migrate -path internal/context/store/migrations -database 'postgres://…' down 1
down restores vector(1536) — the shape migrations 004 and 170 declared — and
truncates first, because ALTER … TYPE against populated rows of the wrong
arity is a hard error. It is conditional: on an install already at 1536 it does
nothing and keeps every row, so stepping the chain on a stock install is safe.
It needs no confirmation GUC. It also gives you nothing back: see the top of this file.
Failure modes and what they mean
refusing to re-dimension … from vector(1536) to vector(768)
You set upsquad.embedding_dim and not upsquad.embedding_redim_confirm. Nothing
was changed. Add the confirmation if you meant it.
this process embeds at 1536 dimensions but agent_memory_embeddings.embedding is vector(768)
FR-7. Either the migration ran and the service configuration did not follow, or
the reverse. The message names both numbers and both remedies.
The most likely instance of this is the dev fixture on a re-dimensioned devbox:
its stamp is dev-hash-embedding-v1 at a fixed 1536, so ENVIRONMENT=development
plus a 768 schema refuses to boot. That is correct — it is a fixture, and it has
no other arity to offer. Point the stack at the real embedder.
Dirty database version -1. Fix and force version.
A migration RAISEd and golang-migrate marked the chain dirty. Migration 210's
failures are atomic — it is a single DO block, so nothing is half-applied —
so the schema is in whichever state it was in before you ran it. Confirm with:
SELECT c.relname, a.atttypmod
FROM pg_attribute a JOIN pg_class c ON c.oid = a.attrelid
WHERE a.attname = 'embedding'
AND c.relname IN ('agent_memory_embeddings', 'context_embeddings');
then migrate … force <version>. See also dev-migrate-dirty-recovery.md.
ERROR: column does not have dimensions
Someone has made one of these columns an unconstrained vector. That column
cannot carry an HNSW index on pgvector 0.8.0. ADR-0031 records this shape as
rejected with that exact error; do not re-propose it.
Verifying an install
-- The arity both tables must agree on.
SELECT c.relname, a.atttypmod AS dimensions
FROM pg_attribute a JOIN pg_class c ON c.oid = a.attrelid
WHERE a.attname = 'embedding'
AND c.relname IN ('agent_memory_embeddings', 'context_embeddings');
-- Both HNSW indexes survive ALTER … TYPE and are rebuilt in place. If either is
-- missing, every ANN query is a sequential scan and nothing says so.
SELECT indexname FROM pg_indexes
WHERE indexname IN ('idx_memory_emb_hnsw', 'idx_embeddings_vector_global');
The boot line each service emits (FR-10) carries provider, model,
dimensions, endpoint_class and fixture — that line and the query above are
the two halves of "is this install consistent".