Skip to main content

Runbook — changing the embedding dimensionality of an install

Owner: backend-sme · Applies to: every install past migration 210 Refs: LLD #2714 §D1 · PRD #996 FR-6/FR-7 · ADR-0031 · issue 2727


The one thing to read before anything else​

Schema rollback is real and takes seconds. Data rollback does not exist and never did. Reversing the schema means re-embedding again. There is no cross-arity projection that preserves meaning, so nothing anywhere in this system can give you your old vectors back once you have confirmed the change. The sources — rag_chunks and agent_memory — are never touched by any step below, which is what makes a re-embed a recovery rather than a loss.

Founder decision FD-4 (PRD #996) says a re-embed is always acceptable. This runbook is written on that basis.


What decides the arity​

One number per install (FD-3). Three things must agree, and the system now refuses to run when they do not:

FactWhere it livesWhat disagreement causes
What the embedder producesprobed from the live endpoint at boot (Dimensions())—
What the schema can holdcontext_embeddings.embedding, agent_memory_embeddings.embeddingFR-7 refuses to boot, naming both numbers
What the migration targetsthe upsquad.embedding_dim connection GUCthe migration no-ops or refuses; never silently half-applies

EMBEDDING_DIM is a claim about a server. The embedder probes the server and the probe wins. If your config says 768 and the endpoint serves 1536, the process fails at construction with ErrServedDimensionMismatch — before FR-7 is even reached.


Procedure — moving an install to a new arity​

Order matters. ALTER … TYPE takes ACCESS EXCLUSIVE, and a row of the old arity present makes it fail rather than corrupt, so the migration cannot be interleaved with a re-embed.

1. Stop the writers​

context-engine, upsquad-memory-mcp, and any running memory-backfill. They will refuse to start again until step 3, by design.

2. Re-dimension the schema​

migrate -path internal/context/store/migrations \
-database 'postgres://…/upsquad?sslmode=disable&options=-c%20upsquad.embedding_dim%3D768%20-c%20upsquad.embedding_redim_confirm%3Dyes' \
up

%20 is a space and %3D is = — the whole options= value rides inside a URL query parameter.

On the dev compose stack, set both in .env instead and run make dev-up once:

EMBEDDING_DIM=768
EMBEDDING_REDIM_CONFIRM=yes

Then remove EMBEDDING_REDIM_CONFIRM again. Leaving it set means the next person who edits EMBEDDING_DIM clears two corpora by starting a stack.

What you will see:

You passedWhat happens
neither GUCno-op, rc=0, every row retained. This is the normal path for every migration harness in the repo.
embedding_dim == the live arityno-op, every row retained — even with the confirmation set.
embedding_dim != live, no confirmationrefusal, naming both arities and the row count. Nothing changes.
bothboth tables truncated and re-typed in one transaction.

3. Point the services at the new embedder​

Set EMBEDDING_PROVIDER / EMBEDDING_BASE_URL / EMBEDDING_MODEL / EMBEDDING_DIM and start them. The FR-7 check runs at boot; if the schema and the embedder disagree you get a message naming both numbers and a non-zero exit, not a pgvector error hours later.

4. Flush the embedding cache (belt and braces)​

redis-cli --scan --pattern 'emb:*' | xargs -r redis-cli del

Postgres TRUNCATE does not reach Redis, and emb:{content_hash}:{model_version} has a 30-day TTL. This step is not load-bearing and you should know why: since LLD #2714 T1/T2 the model identity carries the arity, so a re-dimension necessarily changes model_version and a post-migration read cannot hit a pre-migration entry. The flush is for disk, not for correctness — which is the right way round, because a runbook step is a step someone skips.

5. Re-embed​

Both corpora are empty until you do. Recall does not error in the meantime: retrieval falls back to BM25 + recency and memory recall returns empty.

memory-backfill --batch 200 # agent memory
rag-reembed --batch 200 # rag_chunks / context_embeddings

Both take --dry-run (count only, writes nothing) and --limit N. Run the dry run first: the number it prints is the real backlog under the identity this process resolved, not an estimate — both tools construct their embedder BEFORE counting, precisely so the count is a claim about a known identity.

Both are resumable by re-running them. There is no resume file and no bookmark: a row that has been re-embedded no longer satisfies the candidate predicate, so an interrupted run restarted from the beginning walks only what is left. Interrupt them freely.

Watch the final line of each. converged=true means every candidate the run identified was written; anything else means re-run. A run that ends converged=false with a non-zero failed has skipped individual bad rows and kept going — it will attempt each of them exactly once more per run, so re-run and then investigate whatever still fails.

rag-reembed does not delete rows written under a previous model_version. Recall excludes them already (retrieval/vector.go filters the ANN on the live stamp), and keeping them is what makes a rollback to the previous embedder free rather than a second full re-embed. Migration 211's (chunk_id, model_version) uniqueness bounds that to one row per chunk per model you have actually used. If you are reclaiming disk after a switch you are certain of, that DELETE is a deliberate operator action, not something the tool does unattended.

Check convergence from the outside too. GetSourceStatus reports pending_chunk_count per knowledge source against the live embedder identity — that number going to zero is the operator-visible form of the same fact. Before LLD #2714 T7 it counted any embedding row and therefore read 0 whatever had happened, which is why this step used to look complete the moment it started.


Recovering from an embedding paused degradation (MG-CTX.2)​

rag-reembed is the ONLY recovery, and nothing re-drives those documents for you. A re-dimension is not the only way rag_chunks rows come to exist with no vector. Since LLD #3138 task 5D3 (#3163), an embedding capability that cannot serve — credential absent or expired, endpoint gone, breaker open — puts the capability in the declared embedding paused state and the ingest continues: the chunk rows land, no context_embeddings row does, and recall falls back to BM25 + recency. That is deliberate (the alternative rolled back the whole ingest, taking the source's bookkeeping and audit row with it), and it leaves exactly the condition step 5 above repairs.

What it does not leave is a queue. Configuring a live binding fixes the NEXT ingest and repairs nothing already landed:

  • knowledge ingest is hash-idempotent — an unchanged body hash is a no-op, so a later sync never re-drives a document ingested while paused;
  • the conversation embed-worker acks a degraded run, because the degradation is a success by design.

So after fixing the credential or the binding:

rag-reembed --dry-run # the backlog, under the identity this process resolved
rag-reembed --batch 200 # then the repair; re-runnable, no bookmark

The number to watch is the same one as above — each source's pending_chunk_count — and the fleet-wide view is context_capability_degraded_work_total{capability="embedding",unit="chunk"}, which counts every chunk left unembedded rather than one per incident.


Rolling back​

migrate -path internal/context/store/migrations -database 'postgres://…' down 1

down restores vector(1536) — the shape migrations 004 and 170 declared — and truncates first, because ALTER … TYPE against populated rows of the wrong arity is a hard error. It is conditional: on an install already at 1536 it does nothing and keeps every row, so stepping the chain on a stock install is safe.

It needs no confirmation GUC. It also gives you nothing back: see the top of this file.


Failure modes and what they mean​

refusing to re-dimension … from vector(1536) to vector(768) You set upsquad.embedding_dim and not upsquad.embedding_redim_confirm. Nothing was changed. Add the confirmation if you meant it.

this process embeds at 1536 dimensions but agent_memory_embeddings.embedding is vector(768) FR-7. Either the migration ran and the service configuration did not follow, or the reverse. The message names both numbers and both remedies.

The most likely instance of this is the dev fixture on a re-dimensioned devbox: its stamp is dev-hash-embedding-v1 at a fixed 1536, so ENVIRONMENT=development plus a 768 schema refuses to boot. That is correct — it is a fixture, and it has no other arity to offer. Point the stack at the real embedder.

Dirty database version -1. Fix and force version. A migration RAISEd and golang-migrate marked the chain dirty. Migration 210's failures are atomic — it is a single DO block, so nothing is half-applied — so the schema is in whichever state it was in before you ran it. Confirm with:

SELECT c.relname, a.atttypmod
FROM pg_attribute a JOIN pg_class c ON c.oid = a.attrelid
WHERE a.attname = 'embedding'
AND c.relname IN ('agent_memory_embeddings', 'context_embeddings');

then migrate … force <version>. See also dev-migrate-dirty-recovery.md.

ERROR: column does not have dimensions Someone has made one of these columns an unconstrained vector. That column cannot carry an HNSW index on pgvector 0.8.0. ADR-0031 records this shape as rejected with that exact error; do not re-propose it.


Verifying an install​

-- The arity both tables must agree on.
SELECT c.relname, a.atttypmod AS dimensions
FROM pg_attribute a JOIN pg_class c ON c.oid = a.attrelid
WHERE a.attname = 'embedding'
AND c.relname IN ('agent_memory_embeddings', 'context_embeddings');

-- Both HNSW indexes survive ALTER … TYPE and are rebuilt in place. If either is
-- missing, every ANN query is a sequential scan and nothing says so.
SELECT indexname FROM pg_indexes
WHERE indexname IN ('idx_memory_emb_hnsw', 'idx_embeddings_vector_global');

The boot line each service emits (FR-10) carries provider, model, dimensions, endpoint_class and fixture — that line and the query above are the two halves of "is this install consistent".