Skip to main content

Runbook — installing the self-hosted embedding path

Owner: backend-sme · Applies to: any install running its own embedding endpoint Refs: PRD #996 FR-14 / FR-4 / FR-17 · LLD #2714 §5 T14 · ADR-0031 · milestone 32

This runbook takes an install from "embeddings go to a third party" to "embeddings never leave the perimeter", using configuration files only — no RPC, no UI, no admin screen. That is PRD #996 FR-14's requirement and this page is its acceptance.

A NEW install starts already in the second state. Since FR-9 step 2, EMBEDDING_PROVIDER unset means the bundled in-perimeter embedder, so steps 1 and 2 below are things you verify rather than things you set. Step 3 (the one-time schema move to 768-d) and step 5 (re-embedding an existing corpus) are still real work, and step 3 is still destructive. Read the two blockers below before starting either way.

What every value means: embedding configuration reference. This page is the order to do things in.


Read before you start​

0. Where the #2546 gate stands, since this page is where an operator looks. #2546 — a default install must not egress — is CLOSED (2026-09-03T06:32:09Z) on founder ratification. Both halves are met:

  • No tenant content leaves — measured for the embedding seam: the FR-4 smoke counts ingest and query dials separately and both show google + external == 0 from a network proven routeless. This is a seam property, not a platform property — keep stating it as "no tenant content egresses from the embedding seam, measured", because the measurement is still of one seam.

    The WS chat completion path used to be the named uncovered seam here (#2900, a P1). That caveat is moot: #3585 deleted the /ws/chat seam, its vendor-host variable and its direct dial, so there is no longer a chat egress path in cmd/context-engine to cover. Chat egress now happens in the worker, governed by MG-4.1, and is measured there — not by this runbook.

  • It boots with no egress at all — every blocking fix landed (#2790, #2783, #2792, all closed), and the evidence-quality objection (#2834) is closed too.

Two caveats survived that closure and are deliberately carried here rather than lost with the issue: (1) control C's topology proof comes from the smoke stack — a default install's zero-egress rests on configuration plus the absence of credentials; (2) the artifact's Trivy lane is non-required, so a CRITICAL cannot block a merge. Promote it before any external or on-prem shipping.

1. A fully air-gapped install works. No tokenizer pre-stage is needed. Earlier revisions of this runbook told you to allow egress to openaipublic.blob.core.windows.net on first start, or to pre-warm the tiktoken cache on the host. Both instructions are obsolete — do not follow them.

The cause was that cmd/context-engine and cmd/tools/rag-reembed resolved the cl100k_base BPE vocabulary over HTTPS on first use and treated the failure as fatal, so on a host with no route out they did not start (#2783, #2792). The vocabulary is now compiled into every binary that needs it (internal/context/tiktokenvocab), verified against a SHA-256 pin before use, and the loader fails closed rather than falling back to the network — so a regression surfaces as a named error, never as a silent fetch.

If you already allowlisted openaipublic.blob.core.windows.net for a previous install, remove it: nothing in the platform dials it, and leaving it open is a hole in the perimeter that buys you nothing.

Verified by scripts/embedding-local-egress-smoke.sh assertion H, which asserts no service in the stack has a tokenizer cache staged, on a Docker network with no default route (assertion C).

Read this as "measured on the shipped shape", not as a standing guarantee. The claim names two binaries — cmd/context-engine and cmd/tools/rag-reembed — and it was true when measured. It is not continuously protected: the lane that runs assertion H (sandbox-egress-smoke.yml) is non-required, and its allow-list path filter includes cmd/context-engine/** but omits cmd/tools/rag-reembed/**, so a change to the second binary does not re-run the proof (#2882, open). If you are relying on the air-gap property for an on-prem or external deployment, run the smoke yourself against the build you intend to ship.

2. A fresh install reaches 768-d with EMBEDDING_DIM=768 and nothing else. Migration 210 exempts a zero-row table from the confirmation gate (#2790, closed). A stack that has already embedded something still needs the confirmation, and it is still destructive — see step 3.

3. Step 3 is destructive and there is no data rollback. Moving to the 768-d default clears both embedding tables. The sources — rag_chunks and agent_memory — are never touched, which is what makes a re-embed a recovery rather than a loss. Reversing the schema means re-embedding again.

4. Your users' search queries are in scope, not just your documents. Every search string is embedded, so it goes wherever the embedder points. Making that destination private is the point of this runbook.


What you need​

  • An OpenAI-compatible embedding endpoint reachable from the application containers. Ollama, vLLM, TEI, or a LiteLLM-style facade all speak the shape; there is no adapter and no new provider to configure.
  • The model artifact. For the shipped default this is nomic-embed-text:v1.5, digest-pinned in deployments/embedder/nomic-embed-text.pin.env.
  • 274 MB disk and ~370 MB RAM for the default model, on CPU-only hardware — it needs no GPU. Throughput is 11.5 embeddings/sec and query-path latency 32.3 ms p50 on 8 vCPU. Full numbers and method: footprint measurements.

Procedure​

Step 1 — set nothing​

The default is the answer. With EMBEDDING_PROVIDER unset the process wires the bundled embedder at http://ollama:11434/v1, requests the digest-pinned nomic-embed-text:v1.5, and sends no credential. The model name, the arity and the r= revision come from deployments/embedder/nomic-embed-text.pin.env, which the binary embeds — you do not type them and there is nothing to keep in sync.

What you do need is an OpenAI-compatible endpoint reachable at that name. On compose that is a service called ollama; elsewhere, point the seam at yours:

EMBEDDING_BASE_URL=http://your-inference-host:8080/v1 # optional, must be in-perimeter

A publicly-resolvable host here is refused, and the refusal tells you to set EMBEDDING_PROVIDER=openai if sending text out is what you meant. That is the one deliberate friction in this page: leaving your perimeter stays possible and stops being possible by accident.

Do not set OPENAI_API_KEY for this path. You do not have to remember to: the default provider ignores it entirely, which is the fix for the trap where an ambient key from some other tool ends up on an in-stack wire.

If you are configuring a hosted provider instead, everything in this step inverts — see the configuration reference.

On the tracked dev stack, which chooses Vertex explicitly (#1403), the local path is the docker-compose.local-embedding.yml overlay; skip to make dev-up-local.

Step 2 — verify the endpoint serves what you claim, before touching the schema​

Cheapest possible check, and it is the same request the process makes at boot:

curl -s http://ollama:11434/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{"model":"nomic-embed-text:v1.5","input":"probe"}' \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["model"], len(d["data"][0]["embedding"]))'

Expect nomic-embed-text:v1.5 768. If the model name differs the process will refuse with ErrServedModelMismatch; if the arity differs, with ErrServedDimensionMismatch. Finding out here costs a second; finding out after step 3 costs a re-embed.

Step 3 — move the schema to 768-d (destructive, once)​

The schema ships at 1536 and nomic-embed-text is 768-d, so the columns move once. Migration 210 does both tables in one transaction.

EMBEDDING_DIM=768
EMBEDDING_REDIM_CONFIRM=yes # remove again once the migration has run

Then run migrations. On the dev stack: make dev-up-local.

Leave EMBEDDING_REDIM_CONFIRM unset and the migration refuses, naming both arities, and preserves every row. That refusal is the designed behaviour, not a fault — it converts a destructive act into a decision.

A fresh install does NOT need the confirmation. Migration 210's rung 3 reads IF NOT confirmed AND row_count > 0, so a brand-new stack with zero embedding rows re-dimensions on EMBEDDING_DIM=768 alone — there is nothing for you to decide about (#2790, closed). The confirmation is for a table with rows in it, which is the only case where it protects anything.

Remove EMBEDDING_REDIM_CONFIRM afterwards. Leaving it set arms the destructive path for the next arity change, which is the one you will not be expecting.

For an existing populated install, read changing embedding dimensionality first — it covers the ordering, the rollback answer and the in-flight recall window.

Step 4 — start, and read the boot line​

embedding backend wired — every document this process ingests AND every search
query its users type is sent to the embedding backend you choose
provider=local model=nomic-embed-text:v1.5 dims=768
endpoint_class=private fixture=false
identity=m=nomic-embed-text:v1.5|e=private|d=768|r=970aa74c0a90

provider=local with endpoint_class=private is what a correct default install looks like. endpoint_class=private is the check that matters. It is the whole answer to "does my text leave my perimeter". external or google means it does. fixture=false confirms you are not running pseudo-embeddings. dims is probe-observed from a real response, not your configured claim.

Step 5 — re-embed the existing corpus​

Step 3 cleared the vectors; the sources are intact. Two corpora, two tools, both resumable and both safe to interrupt:

# the RAG corpus
cmd/tools/rag-reembed -batch 200 # -dry-run first to see the candidate count

# agent memories
cmd/tools/memory-backfill

Between the truncate and convergence, retrieval degrades to BM25 + recency and memory recall returns empty. It does not error, which means nothing will page you — watch convergence deliberately.

Watch the tools' logs, not the metrics. memory_reembed_backlog and memory_reembed_processed_total exist, but they are emitted only from these two tool binaries, and neither installs a meter provider — so today they record into a global no-op and you will see nothing on a dashboard. Tracked as #2785 (open). The working observable is the tools' own structured JSON on stderr: rag-reembed: candidates identified, rag-reembed: progress, rag-reembed: done. A second full run is a cheap no-op, so re-running to confirm convergence is the reliable check.

Step 6 — prove it, with a search and not just an ingest​

Ingest one document and run one search. A one-sided check that "no egress occurred" is satisfied by a run in which the query path never embedded at all, which is precisely the bug it would be hiding — so exercise both legs.

scripts/embedding-local-egress-smoke.sh is the mechanised form and it is what CI runs: it counts outbound dials by destination class on the ingest leg and the query leg separately, requires google + external == 0 and private > 0 on both, and carries a control that dials a public address from the routeless network and requires that dial to fail — so the negative half is structural rather than merely observed.


Choosing the quality tier instead​

gte-Qwen2-1.5B-instruct is native-1536 and is available through the same configuration surface with no code change — set EMBEDDING_MODEL, EMBEDDING_DIM=1536, and skip step 3's re-dimension since 1536 is the shipped schema arity.

It costs 5.31× on the synchronous query path — 171.3 ms against 32.3 ms at p50 — and that cost is paid on every search, forever. It also needs 4.07× the disk (1116.7 MB) and 4.23× the RAM (1566.7 MB) at Q4_K_M — a mixed-precision comparison against an f16 default; like-for-like at f16 the disk ratio is 12.97×. Disk and RAM are paid once; query latency recurs.

Two things before you commit to it:

  • You supply the weights, and we do not vouch for them. It is documented and supported, not shipped — which keeps our supply-chain delta at exactly one artifact. Upstream publishes no GGUF, so pick a quantization deliberately: the choice moves disk by 4.7×, and the RAM and latency figures above are Q4_K_M specifically. The artifact is community-produced, unvalidated by us, and NOT digest-pinned under ADR-0029 — unlike the shipped default, which is vendored into the image build and pinned to r=970aa74c0a90. We measured it; we did not endorse it. You own its provenance, its integrity checking and its updates.
  • We publish no retrieval-quality claim for it. Nothing in this milestone measured its recall. It is called the "quality tier" because it is native-1536 and larger — not because it was measured to retrieve better; no dimensionality comparison was performed. If quality is your reason, measure it on your own corpus with scripts/embedding-recall-bench.sh before committing.

Full matrix: configuration reference.


If it refuses to start​

The process fails loudly rather than booting degraded — an explicitly chosen provider that cannot be constructed is a hard failure, because a nil embedder silently substitutes "no search" for "search". Every refusal names a fix that works in the environment that produced it.

Message namesWhat happenedDo
ErrNoProviderConfiguredEMBEDDING_PROVIDER names something unknown; vertex without VERTEX_PROJECT; or the default local pointed at a public EMBEDDING_BASE_URLUnset is not one of these — it is the local default. For the public-URL case: point it in-perimeter, or set EMBEDDING_PROVIDER=openai if you meant it
ErrEndpointUnreachableThe endpoint did not answer the startup probe, or lacks the modelCheck reachability from inside the container; pull the model. Raise EMBEDDING_PROBE_TIMEOUT for a slow cold load
ErrServedModelMismatchThe response's model is not what you configuredAlign EMBEDDING_MODEL with what the server actually serves
ErrServedDimensionMismatchThe endpoint served a different arity than you claimedFix EMBEDDING_DIM, or unset it and let the probe decide
ErrProviderRefusedInEnvA fixture provider (dev) outside a fixture environment. Since #2889 this is the sentinel's only cause — it previously also covered any real provider under ICS_ENABLE_REAL_LLM=false, which no longer affects the embedding seamUnset EMBEDDING_PROVIDER — the default is the keyless in-perimeter configuration
migration 210 refusing, naming two aritiesThe confirmation option was absent and the table has rowsSee step 3. A zero-row table is exempt (#2790), so this refusal means you have a corpus to re-embed
FR-7 boot check refusing, naming two aritiesThe schema and the embedder disagreeThe schema did not move, or moved to the wrong arity. Re-run step 3

A boot line reading fixture=true on a real install is the failure to care about most, because nothing errors: ingest and search both "work" and the corpus is a bag of words. Assert fixture=false and endpoint_class=private in whatever checks your deployment already runs.


Rollback​

Schema rollback is real and takes seconds. Data rollback does not exist and never did. Migration 210's down restores vector(1536), truncating first because ALTER ... TYPE against a populated table of the wrong arity is a hard error. Reversing the schema means re-embedding again under the old model — the sources survive, the vectors do not.

To go back to a hosted embedder, set EMBEDDING_PROVIDER accordingly, reverse the arity if it differs, and re-embed. Nothing about this path is one-way except the vectors themselves — and note that going back is now an explicit act in the config, which is the whole point: it is a line somebody wrote, not a default somebody inherited.