Embedding configuration reference
Status: normative. PRD #996 FR-14 / FR-17 / FR-19 · LLD #2714 §5 T14 · ADR-0031 · milestone 32.
This page is the complete configuration surface for the embedding seam. It is the file/env surface, which PRD #996 FR-14 makes the MVP surface in its own right: a documented config file alone produces a working install, with no RPC and no UI in the loop. There is no admin screen for this and none is required.
To perform an install, follow the self-hosted embedding install runbook. This page is what the values mean. To change the arity of an install that already has vectors in it, see changing embedding dimensionality.
Where your text goes — read this before choosing. Every document this process ingests and every search query its users type is sent to the embedding backend you choose. Queries are the part operators do not expect to leave:
retrieval/service.goandassembly/memory_loader.gocallEmbedon the user's raw search string, and nothing redacts it on the way past. The process prints this same sentence at boot and in every refusal (embedfactory.egressNotice) so the fact is attached to the moment of the choice, not filed in a document. The default —EMBEDDING_PROVIDERunset — is the configuration where nothing leaves.
The provider set
EMBEDDING_PROVIDER is optional, and unset means local: the bundled
in-perimeter embedder, serving the digest-pinned nomic-embed-text:v1.5
artifact this repository ships. That is FR-9 step 2, and it is what closes
#2546's rule that
silence must never resolve to a third party.
This default has had three values, and the sequence is worth knowing because it tells you what an older install of ours does:
| Until | Unset meant | Why it changed |
|---|---|---|
| T4 | openai in code, vertex in the shipped compose — two different vendors, one layer apart | Which vendor received a tenant's documents depended on which layer you read, and neither reading was a decision anybody had made |
| T4 → T15 | a refusal. The process exited non-zero naming the choices | That converts the egress into a decision, which is not the same as removing it: an operator who types openai because it is the only name they recognise is where they started |
| now | local | Silence resolves inside your perimeter |
A default is not the defect; a default that names a third party is. The rule is about where silence points, not about whether silence is allowed to point anywhere.
The table below is generated from embedfactory.providers, the single
declaration of the provider set
(internal/context/embedding/embedfactory/factory.go). It is not maintained by
hand: TestConfigReferenceProviderTableMatchesTheMap fails if it drifts from
the map, so a provider cannot ship undocumented and a documented provider cannot
outlive its removal.
EMBEDDING_PROVIDER | Kind | Selectable when ENVIRONMENT is |
|---|---|---|
dev | Fixture — never a deployment option | development, sandbox |
local | Deployment option | any value, including unset |
openai | Deployment option | any value, including unset |
vertex | Deployment option | any value, including unset |
local is the default, and it is a real model
local is the OpenAI-compatible wire shape with three things decided for you,
each of which a default install must know without being told:
| Value | Overridable | |
|---|---|---|
| Endpoint | http://ollama:11434/v1 | Yes — EMBEDDING_BASE_URL, if it is in-perimeter |
| Model / arity / revision | nomic-embed-text:v1.5 · 768 · 970aa74c0a90 | Yes — EMBEDDING_MODEL, EMBEDDING_DIM, EMBEDDING_REVISION |
| Credential | none, ever | No |
Two behaviours are specific to it, and both exist because it is what silence resolves to:
- It refuses an endpoint outside your perimeter. If
EMBEDDING_BASE_URLnames a publicly-resolvable host,localrefuses to construct and tells you to setEMBEDDING_PROVIDER=openaiif that is what you meant. Without this, "provider unset" would quietly mean "send everything wherever that variable happens to point" — and that variable gets set by shell profiles, copied.envfiles and inherited manifests. Sending text out of your perimeter stays possible; it stops being possible by accident. - It sends no
Authorizationheader. Not "an empty one": the credential path is not taken at all. An exportedOPENAI_API_KEYis the normal state of any host that has ever run the hosted arm, and forwarding it to an in-stack Ollama is how a "local" stack comes to carry a third-party credential.
The three artifact values are read from
deployments/embedder/nomic-embed-text.pin.env — the same file compose delivers
through env_file: and the image build verifies against the registry. The
binary embeds it with go:embed rather than restating it, so a digest bump
cannot leave a default install stamping r= with weights it is not running.
local and openai pointed at the same endpoint produce the same identity
stamp, so an install configured the T8 way can delete both lines and take the
default without orphaning a single vector.
openai is the self-hosted provider you configure yourself
Under FR-17 there are exactly two wire shapes, and "self-hosted" is not a
third one. openai means the OpenAI-compatible wire shape —
POST {base_url}/v1/embeddings — which is spoken by hosted OpenAI, Ollama,
vLLM, TEI and any LiteLLM-style facade. Pointing EMBEDDING_BASE_URL at an
in-stack Ollama is the whole of a self-hosted deployment. Adding a new
inference server to the supported set is a configuration change, never an
engineering one.
dev is a fixture and is not the local option
EMBEDDING_PROVIDER=dev is a deterministic offline hashing trick with no
model behind it. Its similarity reduces to lexical overlap, so a corpus built
on it is not a corpus — the recall benchmark measures it at 0.350 primary
recall@5 against the local model's 1.000, which is what a bag of words scores.
It exists to make the ingest→embed→retrieve loop exercisable with no credential
and no egress, and it is refused outside the environments named above.
It is not the answer to "how do I run without egress." That answer is the
default, local, which runs a real model inside your perimeter. ADR-0031 A12
rejects the fixture as the local option explicitly, and the recall figures above
are why: the fixture is not a cheaper embedder, it is not an embedder.
Every setting
Read by embedfactory.EnvFromOS() — the one reader of these variables across
every binary that embeds (cmd/context-engine, cmd/upsquad-memory-mcp,
cmd/tools/memory-backfill, cmd/tools/rag-reembed, cmd/tools/recall-bench).
A setting means the same thing in all of them.
Nothing is required
A complete configuration is the empty one. Every variable below has a default, and the defaults compose into a working in-perimeter install.
| Variable | Default | Meaning |
|---|---|---|
EMBEDDING_PROVIDER | local | One of the table above. Unset is the bundled in-perimeter embedder. An unrecognised value is still a refusal — a typo must not silently become the default. |
local (the default)
| Variable | Default | Meaning |
|---|---|---|
EMBEDDING_BASE_URL | http://ollama:11434/v1 | Server root including /v1. Must be in-perimeter — a publicly-resolvable host is refused; use openai to send text out deliberately. |
EMBEDDING_MODEL | nomic-embed-text:v1.5, from the pin | Overriding it also drops the pinned EMBEDDING_DIM and EMBEDDING_REVISION defaults, because they describe that artifact and stamping them onto another model would make r= a lie. |
EMBEDDING_DIM | 768, from the pin | A claim that the startup probe checks against a real response. |
EMBEDDING_REVISION | 970aa74c0a90, from the pin | The r= segment. |
OPENAI_API_KEY | — | Ignored. local sends no credential, whatever is set. |
EMBEDDING_PROBE_TIMEOUT | the embedder's default | Bounds the startup canary probe including its retries. Raise it where a cold Ollama is slow to load. |
openai (the OpenAI-compatible provider you configure yourself)
| Variable | Default | Meaning |
|---|---|---|
EMBEDDING_BASE_URL | (empty ⇒ hosted api.openai.com) | Server root including /v1, e.g. http://ollama:11434/v1. A private/RFC1918 URL is permitted unconditionally in every environment — see Trust class below. |
EMBEDDING_MODEL | the embedder's shipped default | The model to request. Echoed back by Ollama, and the startup probe compares the response's model against this — a mismatch refuses to construct. |
EMBEDDING_DIM | (unset ⇒ whatever the server serves) | A claim about arity that gets checked, not a setting that gets trusted. Unset is correct for Ollama and TEI, where arity is a property of the weights. Sent on the wire only when set. |
OPENAI_API_KEY | (empty) | Required for hosted api.openai.com, optional for anything else. cmd/context-engine may resolve it from Vault instead. |
EMBEDDING_REVISION | (empty ⇒ none) | The pinned digest of the served weights, rendered as the identity's r= segment. Mandatory for private endpoints — see Identity. |
EMBEDDING_PROBE_TIMEOUT | the embedder's default | Bounds the startup canary probe including its retries. |
The Vertex provider
| Variable | Default | Meaning |
|---|---|---|
VERTEX_PROJECT | (none — required) | Missing is a refusal; there is no sensible default for someone else's GCP project. |
VERTEX_LOCATION | us-central1 | |
VERTEX_EMBEDDING_MODEL | gemini-embedding-001 |
Environment gates
| Variable | Default | Meaning |
|---|---|---|
ENVIRONMENT | (unset) | Gates fixture providers and nothing else, matched exactly. Unset is not development and does not license a fixture: an absent value is the absence of a claim, not a claim that this is a dev stack. |
ICS_ENABLE_REAL_LLM was listed here until #2889. It gated non-fixture
providers: when false, every real provider was refused. T15 made an unset
EMBEDDING_PROVIDER resolve to the bundled in-perimeter local embedder, so
the gate no longer guarded the hazard it was written for — "a laptop dialling a
paid API because somebody forgot a variable" — while still refusing local
itself. The embedding seam no longer reads it. The variable still exists and
is still consumed by cmd/context-engine for unrelated concerns (the WS chat
LLM mock arm, the guardrail signing key); it has no effect on which embedder you
get.
The model matrix (FR-19)
Both models are reachable through the configuration above with no code change; the quality tier is a config value, not a build. Adding a third model is done by measuring it and adding a row — never by adding a provider.
| Shipped default | Quality tier — operator-supplied | |
|---|---|---|
| Model | nomic-embed-text:v1.5 (137M, Apache-2.0) | gte-Qwen2-1.5B-instruct (1.78B) |
| Dimensions | 768 | 1536 (native) |
| Supplied by | Vendored into the image build, digest-pinned (FD-1 / ADR-0029) | You. Not shipped — see below |
| Disk | 274.3 MB (f16) | 1116.7 MB at Q4_K_M (4.07× — a mixed-precision comparison; like-for-like at f16 it is 12.97×) |
| RAM at inference | 370.0 MB | 1566.7 MB (4.23×) |
| Bulk ingest | 11.5 embeddings/sec | 2.0 embeddings/sec (5.65× slower) |
| Query-path latency | 32.3 ms p50 / 46.1 ms p95 | 171.3 ms p50 / 235.4 ms p95 (5.31×) |
Measured on 8 vCPU CPU-only hardware, on the same fixture the recall benchmark uses. Method, conditions, raw JSON and the caveats are in the footprint measurements — cite that file in sizing decisions, not this table.
Claim-bound 2 — the Dimensions row is a spec; this table is not a dimensionality result
The 768 and 1536 above sit over four cost rows the 768-d model wins by 4–5×, and that is not a finding about vector width. It is a comparison of two specific artifacts, and the cost gap tracks parameter count — 137M against 1.78B, a 12.99× ratio — not the 2× difference in dimensions. A 1536-d model at 137M parameters would not be predicted by this table, and none was measured.
This milestone produced zero dimensionality evidence, and no claim in either direction is licensed — not "768 is better", not "1536 is better". Both arms of the retrieval benchmark below are 768-d; the only 1536 figures anywhere in this milestone belong to a hashed bag-of-words fixture with no model behind it and to a hosted arm that was never run. Neither is evidence about 1536-d embedding models.
The correct statement is "we did not measure it, and it was measurable" — never "there was nothing to find." Binding per founder ratification of PRD #996 v2.2 (
5521553923), which makes "the 768-vs-1536 dimensionality trade is measured" false.Nothing here is a retrieval-quality claim for the quality tier either — its recall was not measured at all. Larger and native-1536 is not evidence of better retrieval.
The quality tier's cost, stated where you choose it
The quality tier costs 5.31× on the synchronous query path — 171.3 ms against
32.3 ms at p50 — and that cost is paid on every search, forever. Disk and RAM
are paid once at install; query latency recurs. That is the trade, and it is the
reason nomic-embed-text is the default rather than an accident of what was
convenient.
Two things you must know before selecting it:
- There is no canonical artifact, and we do not vouch for the one we measured. Upstream ships no GGUF; every one is a third-party requantization, and the choice moves disk by 4.7× on its own (752 MB at Q2_K to 3559 MB at f16). The numbers above are Q4_K_M. Pick your quantization, then re-run the footprint bench against it — the RAM and latency figures do not transfer. The artifact is community-produced, unvalidated by us, and NOT digest-pinned under ADR-0029, unlike the shipped default. We measured it; we did not endorse it.
- We publish no retrieval-quality claim for it. It is named "quality tier"
because it is native-1536 and larger, not because we have measured it to
retrieve better. Nothing in this milestone measured its recall. If quality is
why you are reaching for it, measure it on your own corpus first —
scripts/embedding-recall-bench.shtakes an--armand a--dim.
Supplying this model yourself is deliberate: shipping it would be a second supply-chain decision of the kind FD-1 made once, and NFR-2 holds the delta at exactly one artifact.
Retrieval quality of the shipped default
On the fixed 40-document benchmark fixture, nomic-embed-text at 768-d placed
the intended answer in the top 5 for 20 of 20 queries.
| metric | local-nomic (shipped) | local-arctic (alternative) | dev fixture (negative control) |
|---|---|---|---|
| primary recall@5 | 1.000 — 0/20 intended answers missed | 0.900 — 2/20 missed | 0.350 |
| graded recall@5 | 0.883 | 0.742 | 0.375 |
| MRR@5 | 0.883 | 0.722 | 0.304 |
| nDCG@5 | 0.879 | 0.720 | 0.297 |
local-arctic is snowflake-arctic-embed:110m — a same-weight-class, same-768-d
self-hosted alternative, measured on the identical corpus, queries and labels.
nomic is at least as good on every one of the 20 queries and strictly better
on nine of them (nDCG, exact sign test p = 0.0039, computed over the 18 queries
where both models found the primary answer, so it does not rest on arctic's two
outright misses). It is not a general ranking of the two models, and no such
claim is made. Full result and caveats:
the recall benchmark.
The stronger reason the default did not move is not in these numbers.
internal/context/embedding/chunker.go targets 564-token chunks plus a 64-token
overlap and keeps code blocks whole to 2048 — it is calibrated to nomic's
2048-token window. A 512-token embedder would truncate ordinary production
chunks silently. The benchmark fixture's documents are far shorter (longest: 104
tokens), so the fixture retires that confound for the fixture only.
Quote the pair, never the graded figure alone. A bare "0.883" reads as "roughly one query in nine gets nothing useful back", which is false here: the entire shortfall is secondary documents, and every primary answer was retrieved. The five imperfect queries each missed only a grade-1 near-miss.
Scope — this is not a platform quality claim. 40 short synthetic documents · 20 single-hop paraphrase queries · one primary answer each · no unanswerable query · vector-only scoring. It is fit for separating two embedders on one corpus and for standing as a regression tripwire. Real traffic arrives full of exact project vocabulary (issue numbers, file paths, ADR ids), is served by the hybrid stack rather than the ANN alone, and runs against far larger and heavily near-duplicated corpora where the absolute number does not transfer.
The 0.5 floor in the harness is a control tripwire, not a quality bar.
PRD #996 NFR-1 sets no numeric threshold — it requires only that the
benchmark runs and that the result ships either way. The floor is calibrated so a
known-bad ranker fails and for nothing else. We have no product bar for
retrieval quality, and a real one would be relative — "local within N points
of hosted on the same corpus".
The hosted 1536-d arm is not measured. No provider credential and no egress route existed on the measurement host. No comparison between local and hosted is made or implied anywhere on this page. #2795 — the request to provision it — is closed as WAIVED, not open. What was declined is a new vendor relationship and a new secret surface, not the ~60 calls; the "fraction of a cent" framing understated the ask. Treat this bound as permanent, not as a hole about to be filled.
Full detail, per-query artifacts and the bindings: recall benchmark measurements.
Trust class — why a private URL needs no override
ADR-0031 D4 keys the SSRF posture on the provenance of the URL, not on the environment name and not on the deployment shape.
| URL provenance | Trust class | Guard |
|---|---|---|
| Process configuration — env or file, set by whoever runs the process. The only source on this seam. | Trusted, same class as DATABASE_URL | None. A private endpoint is permitted unconditionally, in every environment. |
| Tenant-supplied and stored | Untrusted input | The full ADR-0025 egress machinery — and this row does not exist on this seam. |
So an on-prem install with an RFC1918 embedding endpoint boots and serves with no override flag, in production, by design. Keying on "is this prod?" would block the intended deployment; that is the inversion ADR-0031 records.
test/lint/embedding_url_provenance_test.go asserts the base URL is sourced only
from process configuration, and reddens the day a DB-sourced or tenant-scoped
path appears — at which point the trust class has changed and the guard is owed.
Model identity, and why r= matters
Every vector is stamped with a four-segment identity, all segments always emitted:
m=<served model> | e=<endpoint class> | d=<dims> | r=<artifact revision>
The shipped default stamps
m=nomic-embed-text:v1.5|e=private|d=768|r=970aa74c0a90.
e=is the bounded endpoint class —none|google|private|external|unknown.privateis the whole answer to "does my text leave my perimeter". The hostname is deliberately never logged or used as a metric label: under FR-17 it is operator-supplied, so it would set our cardinality and our disclosure from someone else's config file.r=is the pinned weights digest, and it is load-bearing. A digest bump under an unchanged tag changes the vector space while leaving model name, endpoint class and arity identical — a silent mixed corpus that the recall filter structurally cannot see.r=is what makes it visible.
Consequence: changing the pin is a corpus event, not a dependency bump. It makes every existing vector foreign to the recall filter and forces a full re-embed of both corpora. That is correct and it is not free; the refresh cadence (ADR-0031 D5.3 — each minor release, at minimum every six months) is affordable only because the re-embed tools exist.
What the process tells you at boot
One line, from every binary that embeds, carrying the FR-11 notice. This is what a default install prints, with nothing configured:
embedding backend wired — every document this process ingests AND every search
query its users type is sent to the embedding backend you choose
provider=local model=nomic-embed-text:v1.5 dims=768
endpoint_class=private fixture=false
identity=m=nomic-embed-text:v1.5|e=private|d=768|r=970aa74c0a90
dims is probe-observed, not the configured claim: the embedder performs one
canary embed at construction and reads the arity off a real response. A server
that serves 1536 when asked for 768 fails there, before any row is written.
The five refusals
Each is a sentinel that callers match by identity, and each names a fix that is actually accepted in the environment that produced it.
| Sentinel | Cause |
|---|---|
ErrNoProviderConfigured | EMBEDDING_PROVIDER names something unknown, or its own required config is missing (vertex with no VERTEX_PROJECT, local with a public EMBEDDING_BASE_URL). Unset is no longer one of these — it is the local default |
ErrProviderRefusedInEnv | A fixture outside a fixture environment, or a real provider while real providers are disabled |
ErrEndpointUnreachable | The endpoint did not answer the startup probe, or does not have the model |
ErrServedModelMismatch | The response's model is not what was configured |
ErrServedDimensionMismatch | The endpoint served a different arity than was claimed |
An explicitly chosen provider that fails to construct is a hard failure, not a degraded boot. A nil embedder silently substitutes "no search" for "search", which is the silence #2546 exists to remove.
Known limitations
Stated here rather than discovered at install time.
An air-gapped default install works. Both halves of
#2546's gate — no
tenant content leaves and it boots with no egress at all — are met, and
#2546 is CLOSED (2026-09-03T06:32:09Z) on founder ratification. FR-9 step 2
shipped in #2737
(d6eecea4): "unset means local" is now true of the provider selection, of the
schema, and of a clean boot on a host with no route out at all.
The exact scope of "no tenant content leaves", because the sentence is broader than the proof. It is measured for the embedding seam — the FR-4 smoke counts ingest and query dials separately and both show
google + external == 0from a network proven routeless. It is not a platform-wide property: the WS chat completion path is a separate seam and is tracked by #2900, an open P1. State this as "no tenant content egresses from the embedding seam, measured" — never unqualified — until #2900 closes.
| Status | |
|---|---|
cmd/context-engine and cmd/tools/rag-reembed previously resolved the cl100k_base tokenizer over HTTPS on first use and treated the failure as fatal. The vocabulary is now compiled into every binary that needs it (internal/context/tiktokenvocab), verified against a SHA-256 pin, and the loader fails closed rather than falling back to the network. No tokenizer pre-stage and no openaipublic.blob.core.windows.net allowlist is needed — if you added one for an earlier install, remove it. Stated as measured on the shipped shape, not as a standing guarantee: the lane that proves it (sandbox-egress-smoke.yml) is non-required and its allow-list path filter omits cmd/tools/rag-reembed/**, the second of the two binaries in the claim — so a change to that binary does not re-run the proof (#2882, open). | #2783, #2792 — CLOSED by PR #2829 (75578032) |
IF NOT confirmed AND row_count > 0: the confirm gate applies to a populated table and is exempt on an empty one, so EMBEDDING_DIM=768 alone brings up a clean stack. The confirmation is still required — and still destructive — for a stack that has embedded something. | #2790 — CLOSED |
The re-embed convergence metrics are not exported. memory_reembed_backlog and memory_reembed_processed_total are emitted only from cmd/tools/rag-reembed and cmd/tools/memory-backfill, and neither binary installs a meter provider — so they record into a global no-op. Watch the tools' own structured progress logs instead. | #2785 — open |
| The hosted 1536-d arm is unmeasured, so no local-vs-hosted comparison exists — and this is permanent, not pending. The measurement was waived, not deferred: the founder declined a new vendor relationship and a new secret surface to answer a question no current decision depends on. The bound accepted in exchange — no local-vs-hosted comparative claim, ever, and say "we did not measure it, and it was measurable" — therefore does not expire. It lifts only if a hosted embedding credential comes to exist for some other reason (Model Gateway or BYOK), in which case run the arm and record it against G5. | #2795 — CLOSED as WAIVED (2026-09-03T06:32:13Z) |