ADR-0031 — Self-hosted embedding: dimension-flexible schema, model identity that names the served model, and the SSRF trust-class inversion
- Status: Proposed — awaiting founder approval on #2723. Not approved, not implemented. The four founder decisions FD-1…FD-4 were already approved on 2026-08-28 on #996; this ADR records them with their rationale and does not reopen them. What is genuinely being approved here is the architect's design under those decisions — principally D4's security-posture inversion and FD-1's supply-chain terms in D5.
- Date: 2026-09-01 (proposed)
- Deciders: Principal Architect (author + technical recommendation). Founder sign-off required — this carries three items from the human-decision escalation list: a new artifact entering the supply chain (
nomic-embed-text, under ADR-0029's terms), a security-relevant inversion of the SSRF posture (D4), and a data-destructive migration path (D1's re-dimension). Product Manager reviews the ADR as an artifact; the founder approves the decision. - Refs: #2723 — task T0, where founder approval is recorded · LLD #2714 v1.1 (this ADR records its D1, D3, D4, D9, D10) · PRD #996 v2.1 (FD-1…FD-4, approved 2026-08-28) · milestone #32 · #2546 — the fail-closed exposure this programme clears; its closure is an explicit founder decision only, never a commit keyword, and nothing in this ADR or any PR under it may carry a closing keyword beside it · #2501, #2604 — the merged recall-side model filters this builds on · ADR-0029 — PII detection inside the trust boundary; the governing supply-chain terms for FD-1 · ADR-0030 D3 / PR #2584 — the endpoint-class vocabulary reused here, never re-minted · ADR-0025 — egress / SSRF trust model, whose machinery D4 declines to add to this seam and says why.
Context
The seam
There is exactly one conceptual seam in the embedding path — "which embedder is this process using, and where does it send text?" — and today it has three doors (cmd/context-engine, cmd/upsquad-memory-mcp, cmd/tools/memory-backfill), a hardcoded destination, a model identity that cannot vary, and a schema that fixes vector arity at 1536 in DDL. PRD #996 v2.1 makes self-hosted inference a first-class deployment shape. That is not a new adapter: under FR-17 self-hosted is OpenAI-compatible plus a private URL, so the work is configurability, identity, and schema — not a third wire protocol.
The founder decisions this ADR records (approved 2026-08-28, not reopened)
| Decision | Consequence carried into this ADR | |
|---|---|---|
| FD-1 | Approved — nomic-embed-text (137M, Apache-2.0, 274 MB) enters the supply chain as the shipped local default, digest-pinned, vendored, image-scanned, under ADR-0029's Option A terms | FR-1/FR-4 ungated. NFR-2's single supply-chain delta becomes an approved delta. The one obligation ADR-0029 left outstanding — a stated refresh cadence — is discharged in D5 |
| FD-2 | (b) first — refuse-to-boot-until-explicit-choice ships before local-by-default; (a) follows | FR-9 is two ordered steps. Step 1 depends on neither FD-1 nor FR-6, so the #2546 exposure is cleared early rather than behind a migration |
| FD-3 | Per-tenant embedder selection OUT of this PRD | One dimensionality per install. This is load-bearing for D1 (no per-tenant schema) and decisive for D4 and D6 (no tenant-supplied URL exists on this seam) |
| FD-4 | Option A — nomic-embed-text at 768-d plus the dimension-flexible schema | FR-6/FR-7 in scope; 768 is the shipped default; 1536 stays reachable via the quality tier |
FD-4's rationale, recorded because it will be re-asked. A native-1536 open-weights model does exist — gte-Qwen2-1.5B-instruct (verified). A v2.0-era analysis claimed none did; that claim is false and was withdrawn in PRD v2.1 rather than left to be discovered. The model was therefore rejected as the default, not overlooked, for two reasons that stand on their own: ~10× the footprint, which raises the on-prem hardware minimum; and, being a 1.5B model, it taxes the synchronous query path on CPU — on every search, forever, not only at ingest. It remains an operator option through FR-17's config surface and is documented in FR-19. Shipping it would be a second FD-1-class supply-chain decision; this ADR does not take it (D5.5).
And a third leg, which is the one that disposes of the "just pick a 1536-d model and skip the migration" reply: our 1536 is a vendor knob, not a target. gemini-embedding-001 is natively 3072-d and is truncated to 1536 by a request parameter (outputDimensionality). The schema currently enshrines one vendor's request parameter as a platform constant. So choosing a native-1536 model would not have avoided the dimension-flexible schema — it would only have deferred it, to the first day any operator wanted a model whose arity is neither 1536 nor truncatable to it. FD-4 Option A pays that cost once, now, deliberately.
Evidence base — measured, not assumed
Everything ruled on below rests on probes run against the exact pinned image pgvector/pgvector:0.8.0-pg16@sha256:c72df104…f410f8eb (docker-compose.dev.yml:49) and the real migrate/migrate:v4.17.1, plus code facts re-pinned on a freshly fetched origin/main @ f56d8651.
| # | Question | Verified result |
|---|---|---|
| V1 | Can HNSW index an unconstrained vector column? | No. Verbatim: ERROR: column does not have dimensions. See D2 |
| V2 | Can the live column dimension be read for a boot check? | Yes. pg_attribute.atttypmod on the vector column is the dimension (measured 768 on a vector(768) column). No parsing of format_type needed |
| V3 | ALTER COLUMN … TYPE vector(768) with 1536-d rows present? | Hard error: expected 768 dimensions, not 1536. The re-dimension cannot silently half-apply — it is structurally fail-loud |
| V4 | Same after clearing the rows? | Succeeds, and the HNSW index survives the type change and is rebuilt in place. On a cleared table the rebuild is free |
| V5 | Can a migration read an operator-configured target dimension? | Yes, through the real golang-migrate image, via connection GUC options=-c%20upsquad.embedding_dim%3D768, read by current_setting('upsquad.embedding_dim', true) inside a DO block, with format()+EXECUTE templating the DDL |
| V6 | The full conditional ladder | Verified in one run: target == live ⇒ no-op, rows preserved; target != live without confirmation ⇒ refuse, rows preserved, message names both numbers; with confirmation ⇒ clear + ALTER + index intact + new arity accepted |
Decision
Self-hosted embedding ships as OpenAI-compatible plus a private URL, on a schema whose vector arity is operator-chosen at install time and changed only by an explicitly confirmed, atomic, both-tables migration; every vector carries a four-segment identity that names the model actually served, including the artifact revision, so that a weights refresh is a visible corpus event rather than a silent split; and the SSRF guard keys on the provenance of the endpoint URL — not on environment name and not on deployment shape — which makes a private endpoint set by process configuration unconditionally permitted, in every environment.
D0 — Read this first: the LLD's decision numbers and this ADR's are not the same numbers
LLD #2714 §2 numbers its design rulings D1…D10. This ADR records only the five in T0's scope, plus two rulings that had no LLD D-number, and it renumbers them contiguously. Three numbers collide — the same label means different things in the two documents. The table is here rather than in a footnote because an implementer arriving from an LLD task reference is exactly the person who will be misled by it.
| LLD #2714 §2 | This ADR | Subject | Note |
|---|---|---|---|
| D1 | D1 | Dimension-flexible schema | Same number, same subject |
| — | D2 | Rejected FR-6 shapes, with V1's error text | New here. In the LLD this sat inside D1's "Rejected shapes" paragraph and §6.1; promoted to its own ruling because T0's acceptance criteria name it |
| D3 | D3 | Model identity, four segments, grandfathering | Same number, same subject |
| D4 | D4 | SSRF keyed on URL provenance | Same number, same subject |
| D10 | D5 | FD-1's ADR-0029 supply-chain obligations (OQ-3/5/6) | COLLIDES. LLD D5 is "the dev stub is a fixture" — a different subject entirely. T8 (nomic-embed-text into the supply chain) depends on "T0 for the terms", and those terms are ADR D5, not LLD D5. An implementer following the LLD's numbering lands on the wrong section |
| D9 | D6 | Embedder stays env/file-configured; not aigateway_providers (OQ-1) | COLLIDES. LLD D6 is "the re-embed, two corpora, two shapes" |
| — | D7 | Migration numbering: a gap is not a free slot | New here, and COLLIDES. LLD D7 is "FR-9's two ordered steps". In the LLD this rule lived in §0's evidence table and T5's renumbering note, not as a D-ruling |
Not recorded in this ADR, and deliberately so — LLD D2 (embedfactory, one seam one guard), D5 (the dev stub as fixture — its rejection is recorded here as alternative A12, but the FR-5 enforcement design stays in the LLD), D6 (the re-embed), D7 (FR-9's two steps) and D8 (the startup probe). These are implementation rulings inside the architect's own remit; T0's scope is D1, D3, D4, D9, D10. Cite ADR numbers when referring to this document and LLD numbers when referring to #2714 — never bare "D5" across the two.
D1 — FR-6/FR-7: dimension-flexible schema. Conditional, GUC-templated, confirm-gated re-dimension. Both tables together, atomically.
One migration containing a DO block per table that:
- reads
upsquad.embedding_dimfrom the connection GUC (V5) — absent ⇒ read the live arity and no-op; - reads the live arity from
pg_attribute.atttypmod(V2); - live == target ⇒ returns, touching nothing. This is where NFR-5 / G7 are guaranteed structurally: an existing Vertex install that keeps 1536 cannot be re-embedded by this migration, because the code path that clears rows is not reachable;
- live != target and
upsquad.embedding_redim_confirm != 'yes'⇒RAISE EXCEPTIONnaming both numbers, the row count, and the fact that sources are retained. Rows preserved (V6); - with confirmation ⇒
TRUNCATEthenEXECUTE format('ALTER TABLE … TYPE vector(%s)'). HNSW survives and rebuilds on an empty table (V4).
Why absent-GUC is a no-op and not a refusal. The refusal is not free: 44 Go files apply migrations, each building its own DATABASE_URL, with no shared helper, plus Makefile:169/:173, both compose stacks, hack/verify-app-rw.sh, and four CI lanes. Every one would need the connection option or CI goes red everywhere. Worse, the refusal would be a second copy of one guard — FR-9's "the operator must choose" already lives in the application, and FR-7's boot check catches the specific mismatch this arm would catch. A second copy of a guard is the exact defect FR-18 exists to remove. The fail-safe property is unchanged: absent input destroys nothing.
Why a confirmation gate and not an unconditional drop. The corpus-survival ruling says re-embed is acceptable; it does not say destroying a populated corpus should happen because someone ran make migrate-up. Data-destructive operations are on the human-decision escalation list. The gate converts a destructive act into a decision — the same move FD-2(b) makes for the boot path — at a cost of one connection option in a runbook. It is not a second founder decision: the founder has approved the destruction; the gate only requires that the operator performing it says so.
Why TRUNCATE and not DELETE. Both tables are FK children (context_embeddings → rag_chunks, agent_memory_embeddings → agent_memory) and nothing references them, so TRUNCATE is permitted, transactional, and does not rewrite or bloat a large corpus. It is owner-only and so bypasses RLS — correct here, because under FD-3 a re-dimension is install-wide by construction.
Both tables move together, in one transaction. Non-negotiable, for four reasons: FD-3 gives one dimensionality per install and there is one EMBEDDING_DIM, so a per-table arity is a state with no way to configure it; the same embedder variable feeds both corpora, so two arities would demand two embedders, which FD-3 forbids; the FR-7 boot check would have two answers and would have to pick one, which is how a silent skip is born; and a partial apply leaves one corpus recallable and the other quietly empty — the SILENT-DROP shape. Schema moves atomically; the two re-embeds converge independently.
down must TRUNCATE first. down restores vector(1536). V3 proves ALTER … TYPE against a populated table of the wrong arity is a hard error, so a down that does not clear first fails on precisely the installs where the migration mattered. Say the rollback answer plainly rather than leaving it to be inferred at 2am: schema rollback is real and takes seconds; data rollback does not exist and never did, by FD-4 — reversing the schema means re-embedding again.
The Redis cache survives the TRUNCATE. emb:{content_hash}:{model_version} has a 30-day TTL and Postgres TRUNCATE does not touch it. That is harmless because the key varies with model identity — except for arity-blind stamps, which is exactly why D3's grandfathering is keyed on (provider, arity). The invariant lives in D3; a runbook cache flush is belt-and-braces only, because a runbook step is a step someone skips.
D2 — The rejected shapes for FR-6, with the evidence, so the obvious simplification is not re-proposed
PRD FR-6 offers three shapes and correctly warns the architect not to read the list as a shortlist. The first one cannot work, and it is recorded here with its error text because it is the intuitive answer and someone will propose it again.
Rejected — unconstrained vector with per-index dimensionality. Eliminated by measurement, not by argument. On our pinned pgvector 0.8.0:
postgres=# CREATE TABLE t_unconstrained (id int, emb vector);
CREATE TABLE
postgres=# CREATE INDEX ON t_unconstrained USING hnsw (emb vector_cosine_ops);
ERROR: column does not have dimensions
Why it looks viable on paper, which is the part worth recording. Mixed-arity rows insert perfectly happily into an unconstrained column — measured, two rows at vector_dims 3 and 4 coexisting in one table — so every experiment short of building the index reports success. The failure is at index time, not at insert time. A control on the same server confirms the contrast is real and not an environment artefact: CREATE TABLE t_768 (id int, emb vector(768)) followed by the identical CREATE INDEX … USING hnsw returns CREATE INDEX. An unconstrained column would therefore give us a corpus that ingests and cannot be searched with an index — the worst available outcome, discovered late.
Rejected — per-dimension sibling tables. Multiplies every one of the four recall queries, both RLS policies, and the classregistry entries by the number of live arities, to buy a simultaneity that FD-3 deleted. Paying structural cost for a capability the founder removed from scope is backwards.
Rejected — application-side ALTER at boot. Puts DDL on the service's startup path, which races N replicas.
Adopted — install-time templated DDL (the third PRD shape), per D1.
D3 — FR-3 + EC-12: model identity must name the served model. Four segments, all emitted unconditionally, legacy stamps grandfathered on (provider, arity).
Mirror extraction/provenance.go exactly — same vocabulary, same k=v|k=v shape, same bounded-metric-label discipline. Do not mint a second vocabulary.
m=<served model, as reported> | e=<endpoint class> | d=<dims> | r=<artifact revision>
m— taken from the responsemodelfield, which is discarded entirely today. Requested-vs-served is the whole point.e—none|google|private|external|unknown, from ADR-0030 D3's merged vocabulary.classifyEndpointis currently unexported insideextraction; it must be lifted to a shared package and reused, never re-implemented. Mutation property: delete theprivatearm ⇒ both the extraction tests and the embedding tests redden.d— dimensionality. Two models can share a name across a re-dimension.r— artifact revision, and this segment is why the FD-1 refresh cadence is safe to have at all. A digest bump ofnomic-embed-textunder an unchanged tag changes the vector space while leavingm,eanddidentical — a silent mixed corpus that FR-12 structurally cannot see.rcarries the pinned digest prefix for our shipped artifact andnonefor hosted providers whose weights we genuinely cannot know. It is the exact analogue of provenance's content-derivedp=prompt revision.- All four segments are emitted unconditionally (IDENTITY-DIMENSION discipline): a positional value is never ambiguous, and a fifth segment appends rather than re-interpreting.
The LiteLLM "*"-alias case, honestly. If the server echoes the requested name we learn nothing new from m — and that is precisely what e=private is for. The stamp then reads "this model name was requested and something private answered", which is true, and is exactly the claim ADR-0030 D3 built the vocabulary to make.
EC-12 ruling: refuse, do not merely log. A wrong stamp is a corpus-wide misattribution written per-vector, and — as with ADR-0029's redaction incident — no column holds the truth that would let it be repaired. "Log loudly" is the failure mode this repo has already paid for. Enforcement sits at the startup probe, not at boot-before-any-request.
Grandfathering, and why it exists. Changing the rendering for the three shipped providers would change every embed_model / model_version value, making every existing row foreign to the FR-12 predicate — a forced full re-embed on upgrade, which directly violates NFR-5 and G7. So Identity carries a legacy field; for openai, vertex and dev it renders byte-identically to today's strings (text-embedding-3-small, vertex-gemini-embedding-001-1536, dev-hash-embedding-v1), each pinned by an exact-value test. New identities render structured. A closed, reviewable, three-entry map — not drift.
Grandfathering is keyed on (provider, arity), never on provider alone. State the reason, because the simplification back to provider-only is one line and looks like tidying. The openai legacy stamp is text-embedding-3-small with no arity segment, and this programme newly makes dimensions configurable — which OpenAI honours, serving the same model name at 768 or 1536. Provider-keyed grandfathering would therefore hand two genuinely different vector spaces one identical stamp and one identical cache key: stale 1536-d vectors served into a vector(768) corpus, and an embed_model value FR-12's filter cannot separate. That is the precise IDENTITY-DIMENSION defect this whole ruling exists to remove, reintroduced on one provider by the map meant to protect compatibility. Each legacy entry is therefore pinned at 1536 only; any other arity renders structured. Every existing install is 1536 by construction — the schema forced it — so NFR-5/G7 is untouched while the hole closes. This is one rule, not three exceptions, and it is what makes D1's Redis-cache note an invariant rather than a runbook step.
Metric labels use Identity.MetricLabel(), collapsing any model outside a source-declared registered set to unregistered, exactly as provenance.go does. NFR-4 requires cardinality bounded by an enum and never by an operator-supplied hostname — and under FR-17 the hostname and the model name are operator-supplied by design, so this matters more here than it did for extraction.
D4 — FR-15: the SSRF guard keys on the provenance of the URL. Not on environment name, not on deployment shape.
This is the security-relevant inversion, and it is stated in both directions so that neither has to be inferred.
What v1.1's NFR-1 said: RFC1918 / link-local / loopback blocked, allow_private_endpoints=true as the opt-in, default false in prod. Under on-prem, a private endpoint IS the expected production configuration. A guard keyed on "is this prod?" blocks the intended deployment — it treats the product as the threat.
Deployment shape is the symptom; URL provenance is the invariant. Keying on shape merely requires a shape flag, and a flag that defaults wrong reproduces the ENVIRONMENT == "production" bug in a new costume.
| URL provenance | Trust class | Guard |
|---|---|---|
| Process configuration — env var or config file, set by whoever runs the process. The only source that exists on this seam under FD-3 | Trusted. Same class as DATABASE_URL, REDIS_URL, AI_GATEWAY_LOCAL_BASE_URL | None. A private endpoint is permitted unconditionally, in every environment, with no override flag — including production. Someone who can set this env var can already set the database URL; there is no privilege boundary left to cross |
Tenant-supplied and stored — an mcp_servers.url, a future per-tenant embedder row, anything read back out of the database | Untrusted input. This is what SSRF means | The full ADR-0025 machinery, unchanged and undiminished: internal/mcp/egress always-deny CIDRs; never-overridable metadata and loopback ranges; approval-scoped allow_private and pinned_cidr; re-validation at dial time |
Both rows are live claims. The second row is not hypothetical politeness — the SSRF concern remains fully live for the hosted/SaaS shape, where a tenant-supplied URL is untrusted input and every ADR-0025 control continues to apply exactly as it does today. This ADR removes nothing from that path. What it says is that the embedding seam has no row-2 source: under FD-3 there is no tenant-supplied embedder URL, so the correct amount of SSRF guarding to add to the embedding path today is none — and adding some would block the intended deployment while defending against an input that cannot occur.
What is owed instead is a tripwire, not a guard. A test asserting the embedder base URL is sourced only from process configuration, which reddens the day a tenant-scoped or DB-sourced path to that value appears. The test must assert on the source of the value, not on the value: a test that merely accepts a private IP would stay green after the trust class changed, which is the failure mode that makes tripwires worthless. This tripwire is also the mechanism that makes D6 self-enforcing rather than a promise.
D3's e= segment is D4's compensating detective control, and the pairing is deliberate. Row 1 is a provenance claim, not a destination claim — it says the URL was set by someone who could already set DATABASE_URL, and it therefore permits any endpoint that person configures, including a public one. An operator who points the embedder at a third-party hosted endpoint is doing something row 1 allows unconditionally, and no preventive control here will stop them; that is the direct consequence of declining to guard on destination. What makes that acceptable rather than merely permitted is that it is not invisible: D3 stamps e= on every vector, from the shared ADR-0030 vocabulary, so a public destination renders e=external (or e=google) on the corpus itself and on the bounded metric label — per-vector, at write time, not in a log that rotates.
So the two rulings divide the work: D4 declines the preventive control on a reasoned trust-class argument; D3 supplies the detective one. An operator-configured egress to a third party is permitted, attributable, and countable after the fact — which is exactly the property #2546 exists to obtain, and the reason "trusted provenance" does not degrade into "unobserved". A design that took D4 without D3 would be strictly weaker than either, and the two should not be separated.
D5 — FD-1's supply-chain obligations, discharged under ADR-0029's terms (OQ-3, OQ-5, OQ-6)
ADR-0029 §D1 set the terms for a model artifact entering our supply chain: "pinned by digest, vendored into the image build, and scanned by the existing Container Image Scan gate", and separately noted the outstanding maintenance obligation — "digest pin, image size, scan findings, and a refresh cadence for the weights". FD-1 approved nomic-embed-text under those terms. They are discharged here.
- Digest pin. The artifact is pinned by digest, never by tag.
nomic-embed-text:latestis not a version. Ther=segment of D3's identity is sourced from the pin, never hand-typed — a hand-typed revision is a second copy of a fact, and it will drift from the thing it claims to describe. - In-image vendoring — not pull-on-first-boot (OQ-5). Per FD-1's "vendored into the image build". Two reasons, and the second is decisive: pull-on-boot makes a missing artifact a runtime failure for exactly the air-gapped buyer this GTM targets; and a pulled artifact is outside the Container Image Scan gate, so ADR-0029's terms would be formally unmet while appearing satisfied. One honest deviation, stated rather than hidden: artifact delivery differs by shape — the dev stack pulls into in-stack Ollama, on-prem vendors in-image — while the code path is identical. The zero-egress goal measures the code path.
- Refresh cadence (OQ-3) — reviewed at each minor platform release, and at minimum every six months.
- A digest bump is a corpus event. This is the obligation most likely to be forgotten, so it is stated as a rule and not as a note. Bumping the digest changes
r=(D3), which makes every existing vector foreign to the FR-12 predicate, which triggers exclusion from recall and a full re-embed of both corpora. It is not a dependency bump. It is not routine maintenance. It costs a re-embed, and the cadence in (3) is affordable only because the benchmark harness and both re-embed tools exist to make it so. A refresh performed without that understanding produces a silently mixed corpus ifr=is ever omitted, or a surprise install-wide re-embed if it is not — and the second is the correct behaviour, which is why it must be expected rather than discovered. - The quality tier is operator-supplied (OQ-6).
gte-Qwen2-1.5B-instructis documented and supported, not shipped. Shipping it would be a second FD-1-class founder decision, and this ADR does not take it — which keeps NFR-2's supply-chain delta at exactly one. Its costs are documented where operators choose, not only here: ~10× footprint, and a query-path latency cost paid on every search, forever.
D6 — OQ-1: the embedder stays env/file-configured. It does not join aigateway_providers.
Four reasons, the last decisive:
aigateway_providersis tenant/BYOK-scoped, and joining it re-imports the per-tenant dimension FD-3 explicitly removed.- FR-14 requires a file-only install surface, with no RPC and no UI in the loop.
memory-backfillandupsquad-memory-mcphave no gateway registry client, and would need one purely to construct an embedder.- A DB-sourced
base_urlis stored data, not process configuration — which moves it from row 1 to row 2 of D4's table, and drags the entire ADR-0025 SSRF apparatus onto a seam that demonstrably does not need it.
Revisit only when per-tenant selection returns as its own PRD — at which point D4's tripwire fires, by design. That is the intended coupling: the mechanism that keeps this decision honest is the same test that detects its reversal.
D7 — Migration numbering: a gap is not a free slot
Recorded as a rule because this programme learned it the expensive way, and because it is an instance of the class this repo keeps removing — a no-op that reports success.
The facts, re-derived on a freshly fetched origin/main @ f56d8651: the ceiling is 207 (207_gate_iff_not_allow_widen). 202 and 206 are gaps — no file exists at either number.
The rule: never fill a gap. golang-migrate keeps a single version row and applies only versions strictly greater than it. A migration authored at 206 would therefore be silently skipped on every database already at 207 — every dev stack, beta-dev, and every CI lane that has migrated since. It would apply cleanly on a fresh database and report success everywhere, so CI would be green and the defect would exist only on the installs that already have data. A dimension migration that silently does not run on exactly the installs that already exist is the SILENT-DROP shape, wearing the costume of a tidy number.
Two corollaries bind every migration authored under this programme and after it:
- Always author above the ceiling, never into a gap, even when the gap is visually adjacent and looks reserved.
- Re-derive the ceiling at file time, from a fresh fetch — not from a design document and not from a working tree that may have moved. This repo lands migrations daily. (This ADR's own predecessor LLD recorded a ceiling of 205 in v1.0 and was wrong, because the shared checkout moved underneath it mid-authoring; the durable output is the rule, not the incident.)
- Each migration carries an authored-ceiling test pinning the number it was written against, so the assumption is checked rather than remembered.
Consequences
Positive.
- On-prem becomes a real deployment shape rather than a hosted install with the network turned off. A private inference endpoint is permitted unconditionally by D4, an operator-chosen arity is accepted by D1, and no
.sqlfile has to be hand-edited. - The default install stops egressing tenant content to a third party — the exposure #2546 names. Under FD-2 this lands in two ordered steps, and step 1 depends on neither the artifact nor the migration, so it clears early.
- A weights refresh becomes a visible corpus event instead of a silent quality regression, because
r=makes it legible to FR-12's existing exclusion predicate (D3, D5.4). This is what makes the six-month cadence safe to have. - The re-dimension cannot silently half-apply. V3 makes it structurally fail-loud; D1's atomicity makes a partial apply across the two tables impossible.
- NFR-5 / G7 is guaranteed structurally, not by care. An existing 1536 install that changes nothing cannot reach the code path that clears rows (D1.3), and its stamps render byte-identically (D3).
- A shipped, invisible hole closes:
(provider, arity)grandfathering removes a cache-key and stamp collision that a configurabledimensionsparameter would otherwise create onopenai— same model name, two vector spaces, one identity. - The endpoint-class vocabulary and the SSRF machinery are reused, not rebuilt — ADR-0030 D3's enum is lifted and shared; ADR-0025's controls are left exactly as they are.
Negative / accepted cost.
- A new artifact in the supply chain, with a permanent maintenance obligation: digest pin, image size, scan findings, refresh cadence (D5). ADR-0029 named this cost when it took the same shape of decision; it is real ongoing work that a managed API would have absorbed.
- A digest bump costs a full re-embed of both corpora (D5.4). Bounded and expected, but it is not free and must not be scheduled as though it were.
- A dimensionality change is destructive to derived data, by founder decision. Sources are retained; vectors are not. Data rollback does not exist — reversing the schema means re-embedding again (D1).
- A recall gap during the re-embed window. Between
TRUNCATEand convergence both corpora return zero vector hits; retrieval degrades to BM25 + recency rather than erroring. Note the observability trap this creates: the foreign-model exclusion counter reads zero through this entire window, because the rows are gone, not foreign-stamped. A distinct backlog observable is required — the exclusion counter alone would report health while the corpus is empty. - One extra connection option in the re-dimension runbook (the confirmation GUC). Deliberate: it is what converts a destructive act into a decision.
- The
devfixture cannot write on a 768 install — its stamp is 1536-bound. Correct, since it is a fixture and not a deployment option, butENVIRONMENT=development+ 768 is the standard devbox configuration, so the failure must be loud and named, not a raw pgvector arity error someone has to decode at the call site.
Security.
- No new egress — this decision removes egress. No new SaaS, no new secret surface: the local endpoint is keyless.
- D4 is a genuine relaxation on one axis and must be read as one. A private-IP embedding endpoint set by process configuration is permitted in production with no override flag. The justification is a trust-class argument, not a convenience one: that value is in the same class as
DATABASE_URL, and an attacker who can set it has already won more than SSRF would buy them. The relaxation is scoped to row 1 of D4's table only. - Nothing is relaxed for tenant-supplied URLs. Row 2 keeps the complete ADR-0025 apparatus. The hosted/SaaS shape's SSRF posture is unchanged by this ADR.
- The premise is named so it can be re-examined rather than rediscovered — in the habit ADR-0030 records after ADR-0029's premises proved narrower than they read three times. D4 row 1 holds because FD-3 puts no tenant-supplied URL on this seam. If per-tenant embedder selection ever returns, row 1's reasoning stops applying to the new source and row 2 governs it. The tripwire in D4 exists precisely to make that transition loud, and D6 is the decision that keeps the seam on the right side of it.
- EC-12 refusal is a security property, not just correctness. An endpoint that serves a different model than requested produces vectors stamped with a lie, and no column holds the truth that would allow repair.
Metrics. Endpoint-class cardinality stays bounded by ADR-0030 D3's enum; an operator-supplied model name outside the registered set collapses to unregistered on the metric while the stored row keeps its exact value. Under FR-17 the hostname and model name are operator-supplied by design, so an unbounded label here would be a cardinality incident rather than a theoretical risk.
Alternatives rejected
A1 — Unconstrained vector column with per-index dimensionality. The PRD's first-listed shape and the obvious simplification. Eliminated by measurement, not preference: ERROR: column does not have dimensions on our pinned pgvector 0.8.0. Recorded with its error text and with the reason it survives casual testing — mixed-arity rows insert fine, so only building the index reveals it. See D2.
A2 — Per-dimension sibling tables. Multiplies four recall queries, two RLS policies and the classregistry entries by the number of live arities, to buy simultaneous dimensionalities — which FD-3 deleted from scope. Structural cost for a removed capability.
A3 — Application-side ALTER at boot. Puts DDL on the service startup path and races N replicas.
A4 — Unconditional drop-and-recreate with no confirmation gate. The corpus-survival ruling makes re-embed acceptable; it does not make silent destruction on make migrate-up acceptable. Data-destructive operations are on the human-decision escalation list. Rejected at a cost of one connection option.
A5 — Hard-error when the GUC is absent. The v1.0 position, on the argument that refusal is free. It is not free: 44 migration harnesses with no shared helper, plus two compose stacks and four CI lanes, would each need the connection option. And it would be a second copy of a guard that already exists in the application and at the FR-7 boot check — the exact defect FR-18 exists to remove. Rejected in favour of no-op, which preserves the fail-safe property (absent input destroys nothing) without duplicating a control.
A6 — Provider-keyed grandfathering. Simpler, one map key instead of two, and wrong. With dimensions configurable, OpenAI serves the same model name at 768 and at 1536; a provider-keyed map gives those two vector spaces one identical stamp and one identical cache key. Rejected because it reintroduces, on one provider, the precise defect D3 exists to remove. Recorded here because the simplification back to provider-only is a one-line diff that reads like tidying.
A7 — A data migration rewriting existing stamps old→new, instead of grandfathering. Costs a full MVCC row rewrite of both corpora, and darks recall install-wide if the mapping misses a value.
A8 — Keying the SSRF guard on deployment shape (an on-prem flag). Closer to right than keying on environment name, and still wrong: it requires a shape flag, and a flag that defaults wrong reproduces the ENVIRONMENT == "production" bug in a new costume. Provenance is the invariant; shape is the symptom.
A9 — Registering the embedder in aigateway_providers. Rejected on four grounds (D6), the decisive one being that a DB-sourced base_url is stored data and silently crosses from row 1 to row 2 of D4's trust table — dragging the full SSRF apparatus onto a seam that does not need it, in order to defend against an input FD-3 does not permit to exist.
A10 — Shipping gte-Qwen2-1.5B-instruct as the default. Rejected explicitly, not overlooked — and the v2.0-era claim that no native-1536 open-weights model exists is false and withdrawn. It is rejected for ~10× footprint (raising the on-prem hardware minimum) and for taxing the synchronous query path on CPU, on every search, forever. Available as a documented operator option; shipping it would be a second FD-1-class supply-chain decision.
A11 — Pull the model artifact on first boot. Smaller image, but it makes a missing artifact a runtime failure for the air-gapped buyer this GTM targets, and it puts the artifact outside the Container Image Scan gate, leaving ADR-0029's terms unmet while appearing satisfied.
A12 — Use the existing dev hash stub as the local embedder. This is the most likely wrong turn on this entire PRD, and PRD FR-5 requires this ADR to name it.
DevEmbedder (internal/context/embedding/dev_embedder.go) is a hashing-trick bag-of-words vector, not an embedding model. It produces 1536-d vectors specifically so that it fits the existing column, and its similarity reduces to lexical overlap with no semantic content whatsoever.
Why it is the trap, stated plainly, because every surface property recommends it. It is offline. It is zero-egress. It is already wired into all three binaries. It needs no migration, no artifact, no supply-chain decision, no benchmark — it satisfies, on paper, nearly every headline requirement in PRD #996 at zero cost. It is precisely what someone reaches for when told "use a local embedder", and it would appear to work: ingest succeeds, search returns ranked results, no test fails, and the zero-egress assertion passes. What is actually lost is recall quality, silently and totally — the corpus retrieves on lexical overlap while presenting as semantic search. There is no error, no counter, and no stamp that distinguishes a degraded result from a good one. That is the SILENT-DROP shape applied to relevance rather than to rows, and it is the reason FD-1 was worth paying for at all: the whole point of taking a supply-chain delta was to get a real model, and this alternative quietly gives that up while claiming the win.
Rejected, and enforced structurally rather than by documentation. The existing refusal outside ENVIRONMENT=development stays, and per FR-18 it must hold in all three binaries — today it holds in only one, which is itself the hole. Beyond the refusal: the provider map marks dev with a Fixture flag; the config-reference table is generated from that field, so the stub cannot appear as a deployment option unless someone flips the flag in the declaration; the boot line renders class=fixture; and a lint test asserts the value dev appears in no runbook or deployment manifest as a provider setting. Config validation refuses it outside development because a fixture that is merely documented as a fixture is one hurried operator away from production — the label has to be load-bearing in code, not in prose.
Per FR-5, it is named here in the terms the PRD requires: dev is a test fixture. It is never a deployment option, and it is not "the local option". The local option is FD-1's nomic-embed-text.
Founder decision points
- Ratify D4 — the SSRF trust-class inversion. This is the substantive security item. A private-IP embedding endpoint sourced from process configuration is permitted unconditionally, in every environment including production, with no override flag; the full ADR-0025 machinery is retained unchanged for tenant-supplied stored URLs, and the hosted/SaaS posture does not move. Note the premise this rests on — FD-3 puts no tenant-supplied URL on this seam — and the tripwire (D4) that fires if that ever stops being true.
- Ratify D5 as the discharge of FD-1's ADR-0029 obligations: digest pin, in-image vendoring (OQ-5), six-month minimum refresh cadence (OQ-3), and the operator-supplied quality tier (OQ-6). The item to note explicitly is D5.4: a digest bump is a corpus event, not a dependency bump — it changes
r=, which makes every existing vector foreign and forces a full re-embed of both corpora. Approving the cadence is approving that cost on that cadence. - Note D1's destructive path. A dimensionality change truncates both vector corpora. Sources are retained and the re-embed is tooled and resumable, but data rollback does not exist — reversing the schema means re-embedding again. The confirm gate (
upsquad.embedding_redim_confirm) is an architect-added guard on a destruction the founder has already approved under FD-4; it is not a new decision and is flagged only so the runbook step is expected. - Note the recall gap. Between truncate and convergence, both corpora return zero vector hits and retrieval degrades to BM25 + recency. On a bulk on-prem corpus re-embedded on CPU this window is long.
- Nothing here reopens FD-1…FD-4. They were approved 2026-08-28 and are recorded above with their rationale — including the withdrawn factual claim behind FD-4 — so that the decisions are inherited with their reasoning attached rather than relitigated. If any of them should move, that is a PRD amendment, not an ADR revision.
Not decided here: per-tenant embedder selection (FD-3, out of scope — returns only as its own PRD); shipping the quality-tier artifact (a second FD-1-class decision, D5.5); and the 768-vs-1536 recall-quality number, which is measured by the benchmark and reported regardless of which way it lands.