Recall benchmark — measured results
NFR-1 / G5. LLD #2714 §5 T13 · PRD #996 NFR-1 / FR-19 · issue #2735 · milestone 32.
These are measurements, not prose. T14 owns the operator-facing model matrix; this directory is the evidence it renders, and the four bindings below travel with the numbers. Per-query detail is in the JSON beside this file, so a surprising mean can be read down to the query that caused it without re-embedding anything.
Run of 2026-09-02 (adjudicated labels)
Corpus upsquad-platform-v1: 40 documents, 20 paraphrase queries. Labels
adjudicated by the Product Manager on 2026-09-02 (ruling on #2735): 20/20
primary and 8/8 secondary labels upheld, one addition (q17 gains d36 at
grade 1). Cutoff k = 5.
| arm | model | vector width | ctx | primary recall@5 | graded recall@5 | MRR@5 | nDCG@5 | primary misses | verdict |
|---|---|---|---|---|---|---|---|---|---|
local-nomic | nomic-embed-text:v1.5 (r=970aa74c0a90), private endpoint | 768 — embedding | 2048 | 1.000 | 0.883 | 0.883 | 0.879 | 0 / 20 | PASS |
local-arctic | snowflake-arctic-embed:110m (r=a0161942122c), private endpoint | 768 — embedding | 512 | 0.900 | 0.742 | 0.722 | 0.720 | 2 / 20 | PASS |
dev-lexical | dev-hash-embedding-v1, hashed bag-of-words | 1536 — hash width, not an embedding | — | 0.350 | 0.375 | 0.304 | 0.297 | 13 / 20 | FAIL |
hosted | text-embedding-3-small / gemini-embedding-001 | 1536 — would be; never run | — | not measured — no provider credential and no egress route on the measurement host | — |
Claim-bound 2 — this table is not a dimensionality result, in either direction
Both arms that ran a model are 768-d. This milestone produced zero evidence about embedding dimensionality, and none of the four rows above may be read as bearing on it.
- The
1536in thedev-lexicalrow is the width of a hashed bag-of-words fixture with no model behind it. It is a hash-table artefact of the stub, not an embedding dimension, and its 0.350 measures lexical blindness to paraphrase — the thing the control exists to demonstrate. It is not evidence about 1536-d embedding models. Per ADR-0031 A12,DevEmbedder"produces 1536-d vectors specifically so that it fits the existing column", and its similarity "reduces to lexical overlap with no semantic content whatsoever". So the width is a schema-compatibility choice and the score is explained by the absence of a model — neither is a fact about dimensionality. What the stub would score at another width is unknown and untested, and this page does not guess:devEmbedderDimensionsis a hard-coded constant (internal/context/embedding/dev_embedder.go:35), so the knob does not exist. Nor is width inert for it — the stub hashes modulo that constant (bucket := h.Sum32() % devEmbedderDimensions), so collision rate is a function of the width, which is precisely why no counterfactual may be asserted here in either direction.- The
1536in thehostedrow is a property of a model that was never run. An unmeasured row carries no result at any width.- The two rows that are comparable to each other —
local-nomicandlocal-arctic— are both 768-d, deliberately (see below: a challenger at another width could not be adopted by a re-embed alone). The comparison between them holds width fixed, which is precisely why it says nothing about width.So neither "768 beats 1536" nor "1536 beats 768" is supported here. The honest statement is: we did not measure it, and it was measurable. Never "there was nothing to find." Founder ratification of PRD #996 v2.2,
5521553923(2026-09-03T06:30:40Z), makes "the 768-vs-1536 dimensionality trade is measured" false and binds all downstream and external material.
local-arctic is a self-hosted-against-self-hosted comparison, added because
the hosted arm cannot be run (binding 4). It is not a substitute for it: G5 as
originally written asked for hosted-against-local and this does not answer that
question — see binding 4 for how v2.2 amended G5 and what that did and did not
concede.
The third arm: what was compared, and what the result is
snowflake-arctic-embed:110m was chosen as the closest available counterpart to
the incumbent — 108.89M parameters against 136.73M, BERT against nomic-BERT, F16
against F16, and 768 dimensions against 768. The shared arity is the
deliberate part: a challenger at another dimension could not be adopted by a
re-embed alone, it would need migration 210's rung 3 re-dimension. A 768-d
challenger is one the platform could actually switch to.
The incumbent wins, and the fixture separated the two arms rather than tying
at the ceiling. local-nomic retrieved the intended answer for 20 of 20
queries; local-arctic for 18 of 20, missing q01 and q20 outright (both
scored 0 on every metric — the answer was not in the top 5 at all).
The primary-recall gap is weaker than "2 / 20" reads — it is n = 1
q01 and q20 are the fixture's only reciprocal pair: they share one
document set with the primary and grade-1 roles swapped (q01 → {d01:2, d28:1},
q20 → {d28:2, d01:1}), and no other query in the fixture touches d01 or
d28 at all. Arctic retrieved neither document for either query, so this is
not a d01/d28 confusion — it is one failure region observed through two
windows onto the same two documents. The effective independent sample on
primary recall is therefore n = 1, not n = 2, and a 2/20 primary gap should
not be read as a ranking. (Neither query is atypical in form: 15 and 12 words
against a fixture median of about 15.)
The ranking is nonetheless established — by depth, and without using the two misses
Discard q01 and q20 entirely and compare only the 18 queries where both
models found the primary answer:
| metric | arctic worse | arctic better | tied | exact 2-tailed sign test |
|---|---|---|---|---|
| graded recall@5 | 4 | 0 | 14 | p = 0.125 |
| MRR@5 | 5 | 0 | 13 | p = 0.0625 |
| nDCG@5 | 9 | 0 | 9 | p = 0.0039 |
Across all 20 queries and all four metrics, arctic does not beat nomic on a single query, once. That is the load-bearing evidence. It is significant at p < 0.005 on nDCG, and because it never uses the two misses it is untouched by the n = 1 objection above.
The same conclusion from the other direction — a ceiling argument. If arctic
had performed identically to nomic on the other 18 queries and merely scored
zero on q01 and q20, it would have reached:
| metric | ceiling | observed | shortfall |
|---|---|---|---|
| graded recall@5 | 0.833333 | 0.741667 | 0.091667 |
| MRR@5 | 0.850000 | 0.722500 | 0.127500 |
| nDCG@5 | 0.838156 | 0.720157 | 0.117999 |
It is below that ceiling on all three, so it ranks worse across the remaining eighteen as well. Treating the two misses as contributing zero graded credit is generous to arctic; any grade-1 credit they earn makes the residual deficit larger.
Provenance of this table, disclosed. The PM's ruling on #996 §3-R computed this ceiling as 0.795 / 0.795 / 0.791 and the shortfall as 0.053 / 0.073 / 0.071. The architect found the arithmetic wrong and it is conservatively wrong: the true ceiling is higher, so the true shortfall is larger. The figures above were recomputed from the per-query artifacts a third time, independently, for this document, and agree with the architect to six decimals. The PM's conclusion holds a fortiori; only their numbers move.
What this does and does not license
It supports: "on this fixture, nomic ranks at least as well as arctic on every query and strictly better on nine of them." It does not support a general quality ordering of the two models — one fixture, 40 synthetic documents, one register (Binding 3, which now also carries the production-chunking caveat).
Where it bears on a decision: the fixture might have been ceiling-saturated and unable to distinguish any two competent embedders, since the incumbent already scores the maximum on primary recall. The PM predicted exactly that and recorded the prediction before the run. It was falsified — the graded metrics had roughly 0.12 of real headroom above nomic and the arms separated inside it. So the instrument is more capable than it was credited with being, which cuts both ways: it also means the unmeasured hosted arm was a measurement worth having, not an empty one (#996 §3-R point 6).
No change to the default model is proposed. The incumbent wins on every metric and loses no query anywhere; the challenger's only material advantages — 218 MB against 274 MB on disk, 20% fewer parameters — were not the constraint anyone was up against. The stronger reason is not in these numbers at all, it is the chunker constants in Binding 3.
Context windows differ across the arms, and that was checked before the numbers were read
nomic-embed-text:v1.5 accepts 2048 input tokens; snowflake-arctic-embed:110m
accepts 512. An embedding endpoint does not refuse an over-long input, it
truncates it and returns a well-formed vector for the prefix. A corpus
document above 512 tokens would therefore have been embedded whole by one arm
and clipped by the other, and the resulting gap would have been a measurement of
truncation published as retrieval quality — with the re-embed converging,
coverage complete and every identity stamp correct.
Measured before interpreting any score, with each model's own tokenizer
(Ollama's /api/embed returns prompt_eval_count for the exact bytes the
seeder sends):
| tokens | limit | headroom | |
|---|---|---|---|
longest document (d06) | 104 | 512 | 4.9× |
| longest query | 24 | 512 | 21× |
| inputs at or over 512 tokens | 0 of 60 |
The two tokenizers agreed exactly on all 60 inputs (both are BERT wordpiece), so the count is not itself arm-dependent. Nothing was truncated in either arm and the comparison is clean on this axis.
There is a stronger form of that claim which does not depend on the measurement
at all, and it is the one to quote: the longest input is d06 at 528
characters, so reaching 512 tokens would require more than 0.97 tokens per
character — essentially character-level tokenization, which BERT wordpiece
cannot produce. Even at an absurd 0.5 tokens/character d06 bounds at 264
tokens. Truncation is arithmetically excluded for this fixture, independently
of prompt_eval_count. The measured 104 tokens / 528 characters = 5.08
characters per token is entirely ordinary for English prose.
That argument is confined to this fixture. It does not extend to production
chunk sizes — see Binding 3, item 5. The window is now declared in each pin as
UPSQUAD_EMBED_ARTIFACT_CONTEXT_TOKENS, and
test/lint.TestRecallFixtureFitsTheNarrowestArmContextWindow holds the fixture
inside the narrowest declared window so a future document cannot reintroduce the
asymmetry silently.
The incumbent was re-run on the same host, same session
local-nomic was re-measured immediately after the arctic arm and reproduced
the committed artifact exactly — identical rankings on all 20 queries,
identical means. That is what makes this an apples-to-apples comparison rather
than a new number set beside an old one: same corpus, same queries, same
adjudicated labels, same host, same day. The committed
recall-local-nomic.json is the earlier run and is unchanged.
What this fixture still cannot do
The arms are separated, on depth, at p < 0.005 — that question is settled and re-running this fixture will not add to it. What the fixture cannot do is support a general ranking of two models, and no amount of re-running fixes that. Generalising would need a fixture this one deliberately is not:
- unanswerable queries, so an embedder that returns confident garbage is distinguishable from one that returns nothing;
- near-duplicate documents, which is where 768-d spaces actually differ and where production lives (10³–10⁶ chunks) — 40 mutually-distinct documents is a far easier discrimination;
- multi-hop and multi-answer queries, so recall has a denominator above 1;
- more queries, and more importantly more independent ones — the two primary misses here turned out to be one failure region seen twice, which a 20-query set is small enough to hide;
- production-sized chunks, which this fixture cannot contain without reintroducing the 512-token confound the short documents retire.
Building that is a larger piece of work than T13 and it is not a prerequisite for anything currently blocked.
Binding 1 — publish the pair, never a bare graded recall
Primary recall@5 = 1.000: the incumbent placed the intended answer in the top 5 for all 20 of 20 queries. The entire 0.117 graded shortfall is secondary documents — the five imperfect queries (q01, q11, q15, q17, q20) each missed only a grade-1 near-miss.
A bare "recall@5 = 0.883" reads as "about one question in eight gets nothing
useful back". That is false here, and it is the reading a matrix would ship.
Quote the pair. Result.Summary() renders it that way and a test pins the
ordering, so the graded figure cannot be lifted alone from a run log.
The two are not a flattering restatement — they separate on the control (0.350 primary against 0.375 graded), which is the axis the arms differ on.
The challenger arm is the same story at a different level, and it is the reason the pair is a rule rather than a courtesy to the incumbent: arctic's 0.900 primary against its 0.742 graded. The graded figure would read as "a quarter of questions get nothing useful", when in fact 18 of 20 got their intended answer. Both arms are quoted the same way, including the one that scored lower.
Binding 2 — the 0.5 floor is a control tripwire, not a quality bar
PRD #996 NFR-1 sets no numeric threshold; it requires that the benchmark runs
before the local default flips and that the result ships regardless of which way
it lands. There is nothing here to clear. The floor is calibrated so a known-bad
ranker fails — dev breaches all four conditions — and for nothing else.
We have no product bar for retrieval quality, and this fixture cannot set one. When there is one it should be relative: "local within N points of hosted on the same corpus", which is the form NFR-1 asks for and which needs the hosted arm and a production-scale corpus.
Binding 3 — scope line, adjacent to the number
40 short synthetic documents · 20 single-hop paraphrase queries · one primary answer each · no unanswerable query · vector-only scoring.
Fit for separating two embedders on one corpus and for standing as a regression tripwire. Not a model of real query traffic and not a platform retrieval-quality claim. Four skews, stated because a benchmark that measures the wrong thing while scoring 0.88 is the one that gets cited:
- Every query has exactly one primary answer and none is unanswerable, so the benchmark cannot detect an embedder returning confident garbage for a question the corpus does not answer — a live failure mode for a thin day-one on-prem corpus.
- The paraphrase construction is the inverse of real traffic, which arrives full of exact project vocabulary (issue numbers, file paths, ADR ids). Scoring is vector-only while users are served by the hybrid stack (ANN + BM25 + RRF + recency).
- Forty mutually-distinct synthetic documents is an easier problem than 10³–10⁶ heavily near-duplicated production chunks; chance alone is 0.125. The absolute number does not transfer to production scale.
- One author, one register, one week — queries and documents share a voice.
- The fixture's documents are far shorter than production chunks, so the 512-token confound is retired for THIS FIXTURE ONLY. See below — this is the caveat most likely to be dropped when a number from this page is quoted later, and it is the one that matters most for a 512-token arm.
5, in full: production chunking is calibrated to nomic's window
internal/context/embedding/chunker.go:
targetTokens = 512 // exactly arctic's ENTIRE window
targetMax = 564 // ABOVE arctic's window, routinely
overlapTokens = 64 // prepended on top of that
maxCodeBlockTokens = 2048 // 4x arctic's window; nomic's window to the token
The production chunker is calibrated to nomic's 2048-token window —
maxCodeBlockTokens is nomic's context length exactly. Under a 512-token arm,
ordinary prose chunks at targetMax (564, before the 64-token overlap is
prepended) already exceed the window, and a whole-kept code block or table would
lose up to three quarters of its content. Silently: an over-long input is
truncated, not refused.
So the fixture's cleanliness — longest input 104 tokens, 0 of 60 anywhere near 512 — retires the truncation confound for this fixture and for nothing else. A result measured here licenses no statement about how a 512-token model behaves on production chunk sizes. If anything the chunker constants are an independent and stronger argument against a 512-token embedder than the recall numbers on this page are: the recall gap is a measured tendency, whereas adopting a 512-token model would truncate real content by construction, on the default chunking path, with nothing in the logs to say so.
What the numbers support: "on a fixed 40-document fixture, the shipped local 768-d model placed the intended answer in the top 5 for 20 of 20 paraphrased queries, against 18 of 20 for a same-class 768-d self-hosted alternative and 7 of 20 for a bag-of-words control." Not an absolute quality claim, not an SLA, not a tier or pricing statement, not a comparison against hosted, and not a general statement that one self-hosted model is better than the other — see the two-query caveat above.
Binding 4 — the hosted row: waived, and what stays forbidden
Recorded as "not measured" with its reason — never blank, never a dash. No comparative claim between local and hosted may appear anywhere, explicit or implied: an unmeasured row plus a comparative adjective is a documented claim with nothing behind it.
G5 was amended and the hosted arm was waived — both ratified by the founder on
2026-09-03T06:30:40Z (5521553923,
PRD #996 v2.2). This section previously gated T15 on the hosted arm; that gating
language is retired and is recorded here rather than deleted, because the
earlier revision of this file is cited elsewhere.
- G5 now reads as a cross-model comparison across self-hostable models, not
hosted-against-local. The delivered
local-nomicvslocal-arcticevidence satisfies it as amended. T15 — the local-default flip — shipped (#2737,d6eecea4), and it cleared the gate that governed it. - The reasoning the founder accepted: T15 changes
unset ⇒ refuseintounset ⇒ local. The counterfactual it displaces is refusal, not hosted — nobody is migrated off a hosted embedder because nobody was on one. That is why a hosted comparison was not load-bearing for this flip. - The hosted row is
WAIVED, not pending. #2795 is closed as waived. The founder declined it as unprovisionable rather than merely unprovisioned: no hosted embedding account exists and LiteLLM holds no upstream credential, so the real cost is a new vendor relationship and a new secret surface, not the ~60 calls. Filing it as "~60 calls, a fraction of a cent" understated the ask, and the PM recorded that correction against themselves. - The trigger, for whoever finds this later: if a hosted embedding credential
comes to exist for any other reason (Model Gateway or BYOK being the likely
source), run the arm — one
--arm, one--dim 1536, the provider's own environment. The fixture is now known to discriminate (it separated nomic from arctic at p = 0.0039), so the measurement is known-informative. Record it against G5.
What the waiver bought, and what it permanently costs. The bound below is not a placeholder awaiting the measurement — it was accepted in exchange for the waiver and is therefore permanent unless the trigger above fires:
No comparative claim between local and hosted may appear anywhere, explicit or implied. Say "we did not measure it, and it was measurable" — never "there was nothing to find."
The local-arctic arm does not discharge the hosted question. It is a second
self-hosted model, not a hosted one: same endpoint class (e=private), same
arity, no provider credential involved. It satisfies G5 as amended and it
answers "is there a better self-hosted option we can ship" — a real question,
worth 60 calls. It leaves the hosted row unmeasured, and that row is now closed
as waived rather than left open.
The control arm, and why the first row means anything
dev-lexical is the negative control, not a candidate: a working retrieval
mechanism — well above the 0.125 chance level, with a distinct top hit for most
queries — that is simply blind to paraphrase. It breaches all four floors. Its
job is to make a passing verdict falsifiable.
The same separation is asserted in CI, offline and with no database, by
internal/context/recallbench.TestLexicalArmFailsTheGateOnTheParaphraseQuerySet.
It reddens if the query set ever becomes lexically easy — measured: with every
query replaced by its own answer's text, the lexical arm scores 0.800 / 1.000 /
0.927 and the test fails, naming the cause.
Independently reproduced inside the air-gapped stack
scripts/embedding-local-egress-smoke.sh assertion G runs this benchmark
against the same artifact from a network with no default route (Docker
internal: true, with a control that dials a public address and requires
failure), through retrieval.VectorSearcher — the ANN alone, no BM25, no RRF.
That is the second half of #2546's claim: nothing leaves the perimeter, and
retrieval still works.
Effect of the adjudication, disclosed
The PM made the q17/d36 call from the corpus text before opening any result
artifact, and disclosed that it moves the headline up. Measured effect:
| before | after | delta | |
|---|---|---|---|
| graded recall@5 | 0.875000 | 0.883333 | +0.008333 |
| MRR@5 | 0.883333 | 0.883333 | 0 |
| nDCG@5 | 0.876831 | 0.879468 | +0.002636 |
Control arm unchanged (0.375 / 0.304 / 0.297, 12/20 zero-recall, still fails all four conditions), so arm separation is preserved.
One divergence from the PM's predicted table, and it is arithmetic, not
substance. The ruling predicted nDCG 0.880; the run measured 0.879468. The
estimate rounded twice — 0.877 + 0.003 — where the unrounded values are
0.876831 + 0.002636 = 0.879468. Recall and MRR match the prediction exactly.
q17's nDCG was also recomputed by hand against the retrieved ranking (d36 at
rank 2) and agrees with the engine to six decimal places: 0.878962.
Reproducing
scripts/embedding-recall-bench.sh \
--arm local-nomic --dim 768 \
--env EMBEDDING_PROVIDER=openai \
--env EMBEDDING_BASE_URL=http://127.0.0.1:11434/v1 \
--env-file deployments/embedder/nomic-embed-text.pin.env
# the challenger — same corpus, same queries, same labels, its own scratch DB
scripts/embedding-recall-bench.sh \
--arm local-arctic --dim 768 \
--env EMBEDDING_PROVIDER=openai \
--env EMBEDDING_BASE_URL=http://127.0.0.1:11434/v1 \
--env-file deployments/embedder/snowflake-arctic-embed-110m.pin.env
# the control — expected to FAIL the gate, and the script requires that it does
scripts/embedding-recall-bench.sh \
--arm dev-lexical --dim 1536 --expect-fail \
--env EMBEDDING_PROVIDER=dev --env ENVIRONMENT=development
Adding the third arm needed no change to the harness — an arm is a pin file
plus two settings. Provenance for the challenger is
deployments/embedder/snowflake-arctic-embed-110m.pin.env, derived from the
registry with scripts/embedding-artifact-pin.sh --show and verified against the
live Ollama store with --verify-store before the arm was scored, so the number
above names bytes rather than a tag.
Each arm builds its own scratch database at its own arity and re-embeds the
corpus with cmd/tools/rag-reembed — the shipped FR-8 tool, not a private
embedding path. Re-runnability is what makes ADR-0031 D5.3's refresh cadence
affordable: a weights bump changes r=, which invalidates the corpus, and these
numbers have to be taken again.