Skip to main content

Recall benchmark — measured results

NFR-1 / G5. LLD #2714 §5 T13 · PRD #996 NFR-1 / FR-19 · issue #2735 · milestone 32.

These are measurements, not prose. T14 owns the operator-facing model matrix; this directory is the evidence it renders, and the four bindings below travel with the numbers. Per-query detail is in the JSON beside this file, so a surprising mean can be read down to the query that caused it without re-embedding anything.

Run of 2026-09-02 (adjudicated labels)​

Corpus upsquad-platform-v1: 40 documents, 20 paraphrase queries. Labels adjudicated by the Product Manager on 2026-09-02 (ruling on #2735): 20/20 primary and 8/8 secondary labels upheld, one addition (q17 gains d36 at grade 1). Cutoff k = 5.

armmodelvector widthctxprimary recall@5graded recall@5MRR@5nDCG@5primary missesverdict
local-nomicnomic-embed-text:v1.5 (r=970aa74c0a90), private endpoint768 — embedding20481.0000.8830.8830.8790 / 20PASS
local-arcticsnowflake-arctic-embed:110m (r=a0161942122c), private endpoint768 — embedding5120.9000.7420.7220.7202 / 20PASS
dev-lexicaldev-hash-embedding-v1, hashed bag-of-words1536 — hash width, not an embedding—0.3500.3750.3040.29713 / 20FAIL
hostedtext-embedding-3-small / gemini-embedding-0011536 — would be; never run—not measured — no provider credential and no egress route on the measurement host—

Claim-bound 2 — this table is not a dimensionality result, in either direction​

Both arms that ran a model are 768-d. This milestone produced zero evidence about embedding dimensionality, and none of the four rows above may be read as bearing on it.

  • The 1536 in the dev-lexical row is the width of a hashed bag-of-words fixture with no model behind it. It is a hash-table artefact of the stub, not an embedding dimension, and its 0.350 measures lexical blindness to paraphrase — the thing the control exists to demonstrate. It is not evidence about 1536-d embedding models. Per ADR-0031 A12, DevEmbedder "produces 1536-d vectors specifically so that it fits the existing column", and its similarity "reduces to lexical overlap with no semantic content whatsoever". So the width is a schema-compatibility choice and the score is explained by the absence of a model — neither is a fact about dimensionality. What the stub would score at another width is unknown and untested, and this page does not guess: devEmbedderDimensions is a hard-coded constant (internal/context/embedding/dev_embedder.go:35), so the knob does not exist. Nor is width inert for it — the stub hashes modulo that constant (bucket := h.Sum32() % devEmbedderDimensions), so collision rate is a function of the width, which is precisely why no counterfactual may be asserted here in either direction.
  • The 1536 in the hosted row is a property of a model that was never run. An unmeasured row carries no result at any width.
  • The two rows that are comparable to each other — local-nomic and local-arctic — are both 768-d, deliberately (see below: a challenger at another width could not be adopted by a re-embed alone). The comparison between them holds width fixed, which is precisely why it says nothing about width.

So neither "768 beats 1536" nor "1536 beats 768" is supported here. The honest statement is: we did not measure it, and it was measurable. Never "there was nothing to find." Founder ratification of PRD #996 v2.2, 5521553923 (2026-09-03T06:30:40Z), makes "the 768-vs-1536 dimensionality trade is measured" false and binds all downstream and external material.

local-arctic is a self-hosted-against-self-hosted comparison, added because the hosted arm cannot be run (binding 4). It is not a substitute for it: G5 as originally written asked for hosted-against-local and this does not answer that question — see binding 4 for how v2.2 amended G5 and what that did and did not concede.

The third arm: what was compared, and what the result is​

snowflake-arctic-embed:110m was chosen as the closest available counterpart to the incumbent — 108.89M parameters against 136.73M, BERT against nomic-BERT, F16 against F16, and 768 dimensions against 768. The shared arity is the deliberate part: a challenger at another dimension could not be adopted by a re-embed alone, it would need migration 210's rung 3 re-dimension. A 768-d challenger is one the platform could actually switch to.

The incumbent wins, and the fixture separated the two arms rather than tying at the ceiling. local-nomic retrieved the intended answer for 20 of 20 queries; local-arctic for 18 of 20, missing q01 and q20 outright (both scored 0 on every metric — the answer was not in the top 5 at all).

The primary-recall gap is weaker than "2 / 20" reads — it is n = 1​

q01 and q20 are the fixture's only reciprocal pair: they share one document set with the primary and grade-1 roles swapped (q01 → {d01:2, d28:1}, q20 → {d28:2, d01:1}), and no other query in the fixture touches d01 or d28 at all. Arctic retrieved neither document for either query, so this is not a d01/d28 confusion — it is one failure region observed through two windows onto the same two documents. The effective independent sample on primary recall is therefore n = 1, not n = 2, and a 2/20 primary gap should not be read as a ranking. (Neither query is atypical in form: 15 and 12 words against a fixture median of about 15.)

The ranking is nonetheless established — by depth, and without using the two misses​

Discard q01 and q20 entirely and compare only the 18 queries where both models found the primary answer:

metricarctic worsearctic bettertiedexact 2-tailed sign test
graded recall@54014p = 0.125
MRR@55013p = 0.0625
nDCG@5909p = 0.0039

Across all 20 queries and all four metrics, arctic does not beat nomic on a single query, once. That is the load-bearing evidence. It is significant at p < 0.005 on nDCG, and because it never uses the two misses it is untouched by the n = 1 objection above.

The same conclusion from the other direction — a ceiling argument. If arctic had performed identically to nomic on the other 18 queries and merely scored zero on q01 and q20, it would have reached:

metricceilingobservedshortfall
graded recall@50.8333330.7416670.091667
MRR@50.8500000.7225000.127500
nDCG@50.8381560.7201570.117999

It is below that ceiling on all three, so it ranks worse across the remaining eighteen as well. Treating the two misses as contributing zero graded credit is generous to arctic; any grade-1 credit they earn makes the residual deficit larger.

Provenance of this table, disclosed. The PM's ruling on #996 §3-R computed this ceiling as 0.795 / 0.795 / 0.791 and the shortfall as 0.053 / 0.073 / 0.071. The architect found the arithmetic wrong and it is conservatively wrong: the true ceiling is higher, so the true shortfall is larger. The figures above were recomputed from the per-query artifacts a third time, independently, for this document, and agree with the architect to six decimals. The PM's conclusion holds a fortiori; only their numbers move.

What this does and does not license​

It supports: "on this fixture, nomic ranks at least as well as arctic on every query and strictly better on nine of them." It does not support a general quality ordering of the two models — one fixture, 40 synthetic documents, one register (Binding 3, which now also carries the production-chunking caveat).

Where it bears on a decision: the fixture might have been ceiling-saturated and unable to distinguish any two competent embedders, since the incumbent already scores the maximum on primary recall. The PM predicted exactly that and recorded the prediction before the run. It was falsified — the graded metrics had roughly 0.12 of real headroom above nomic and the arms separated inside it. So the instrument is more capable than it was credited with being, which cuts both ways: it also means the unmeasured hosted arm was a measurement worth having, not an empty one (#996 §3-R point 6).

No change to the default model is proposed. The incumbent wins on every metric and loses no query anywhere; the challenger's only material advantages — 218 MB against 274 MB on disk, 20% fewer parameters — were not the constraint anyone was up against. The stronger reason is not in these numbers at all, it is the chunker constants in Binding 3.

Context windows differ across the arms, and that was checked before the numbers were read​

nomic-embed-text:v1.5 accepts 2048 input tokens; snowflake-arctic-embed:110m accepts 512. An embedding endpoint does not refuse an over-long input, it truncates it and returns a well-formed vector for the prefix. A corpus document above 512 tokens would therefore have been embedded whole by one arm and clipped by the other, and the resulting gap would have been a measurement of truncation published as retrieval quality — with the re-embed converging, coverage complete and every identity stamp correct.

Measured before interpreting any score, with each model's own tokenizer (Ollama's /api/embed returns prompt_eval_count for the exact bytes the seeder sends):

tokenslimitheadroom
longest document (d06)1045124.9×
longest query2451221×
inputs at or over 512 tokens0 of 60

The two tokenizers agreed exactly on all 60 inputs (both are BERT wordpiece), so the count is not itself arm-dependent. Nothing was truncated in either arm and the comparison is clean on this axis.

There is a stronger form of that claim which does not depend on the measurement at all, and it is the one to quote: the longest input is d06 at 528 characters, so reaching 512 tokens would require more than 0.97 tokens per character — essentially character-level tokenization, which BERT wordpiece cannot produce. Even at an absurd 0.5 tokens/character d06 bounds at 264 tokens. Truncation is arithmetically excluded for this fixture, independently of prompt_eval_count. The measured 104 tokens / 528 characters = 5.08 characters per token is entirely ordinary for English prose.

That argument is confined to this fixture. It does not extend to production chunk sizes — see Binding 3, item 5. The window is now declared in each pin as UPSQUAD_EMBED_ARTIFACT_CONTEXT_TOKENS, and test/lint.TestRecallFixtureFitsTheNarrowestArmContextWindow holds the fixture inside the narrowest declared window so a future document cannot reintroduce the asymmetry silently.

The incumbent was re-run on the same host, same session​

local-nomic was re-measured immediately after the arctic arm and reproduced the committed artifact exactly — identical rankings on all 20 queries, identical means. That is what makes this an apples-to-apples comparison rather than a new number set beside an old one: same corpus, same queries, same adjudicated labels, same host, same day. The committed recall-local-nomic.json is the earlier run and is unchanged.

What this fixture still cannot do​

The arms are separated, on depth, at p < 0.005 — that question is settled and re-running this fixture will not add to it. What the fixture cannot do is support a general ranking of two models, and no amount of re-running fixes that. Generalising would need a fixture this one deliberately is not:

  • unanswerable queries, so an embedder that returns confident garbage is distinguishable from one that returns nothing;
  • near-duplicate documents, which is where 768-d spaces actually differ and where production lives (10³–10⁶ chunks) — 40 mutually-distinct documents is a far easier discrimination;
  • multi-hop and multi-answer queries, so recall has a denominator above 1;
  • more queries, and more importantly more independent ones — the two primary misses here turned out to be one failure region seen twice, which a 20-query set is small enough to hide;
  • production-sized chunks, which this fixture cannot contain without reintroducing the 512-token confound the short documents retire.

Building that is a larger piece of work than T13 and it is not a prerequisite for anything currently blocked.

Binding 1 — publish the pair, never a bare graded recall​

Primary recall@5 = 1.000: the incumbent placed the intended answer in the top 5 for all 20 of 20 queries. The entire 0.117 graded shortfall is secondary documents — the five imperfect queries (q01, q11, q15, q17, q20) each missed only a grade-1 near-miss.

A bare "recall@5 = 0.883" reads as "about one question in eight gets nothing useful back". That is false here, and it is the reading a matrix would ship. Quote the pair. Result.Summary() renders it that way and a test pins the ordering, so the graded figure cannot be lifted alone from a run log.

The two are not a flattering restatement — they separate on the control (0.350 primary against 0.375 graded), which is the axis the arms differ on.

The challenger arm is the same story at a different level, and it is the reason the pair is a rule rather than a courtesy to the incumbent: arctic's 0.900 primary against its 0.742 graded. The graded figure would read as "a quarter of questions get nothing useful", when in fact 18 of 20 got their intended answer. Both arms are quoted the same way, including the one that scored lower.

Binding 2 — the 0.5 floor is a control tripwire, not a quality bar​

PRD #996 NFR-1 sets no numeric threshold; it requires that the benchmark runs before the local default flips and that the result ships regardless of which way it lands. There is nothing here to clear. The floor is calibrated so a known-bad ranker fails — dev breaches all four conditions — and for nothing else.

We have no product bar for retrieval quality, and this fixture cannot set one. When there is one it should be relative: "local within N points of hosted on the same corpus", which is the form NFR-1 asks for and which needs the hosted arm and a production-scale corpus.

Binding 3 — scope line, adjacent to the number​

40 short synthetic documents · 20 single-hop paraphrase queries · one primary answer each · no unanswerable query · vector-only scoring.

Fit for separating two embedders on one corpus and for standing as a regression tripwire. Not a model of real query traffic and not a platform retrieval-quality claim. Four skews, stated because a benchmark that measures the wrong thing while scoring 0.88 is the one that gets cited:

  1. Every query has exactly one primary answer and none is unanswerable, so the benchmark cannot detect an embedder returning confident garbage for a question the corpus does not answer — a live failure mode for a thin day-one on-prem corpus.
  2. The paraphrase construction is the inverse of real traffic, which arrives full of exact project vocabulary (issue numbers, file paths, ADR ids). Scoring is vector-only while users are served by the hybrid stack (ANN + BM25 + RRF + recency).
  3. Forty mutually-distinct synthetic documents is an easier problem than 10³–10⁶ heavily near-duplicated production chunks; chance alone is 0.125. The absolute number does not transfer to production scale.
  4. One author, one register, one week — queries and documents share a voice.
  5. The fixture's documents are far shorter than production chunks, so the 512-token confound is retired for THIS FIXTURE ONLY. See below — this is the caveat most likely to be dropped when a number from this page is quoted later, and it is the one that matters most for a 512-token arm.

5, in full: production chunking is calibrated to nomic's window​

internal/context/embedding/chunker.go:

targetTokens = 512 // exactly arctic's ENTIRE window
targetMax = 564 // ABOVE arctic's window, routinely
overlapTokens = 64 // prepended on top of that
maxCodeBlockTokens = 2048 // 4x arctic's window; nomic's window to the token

The production chunker is calibrated to nomic's 2048-token window — maxCodeBlockTokens is nomic's context length exactly. Under a 512-token arm, ordinary prose chunks at targetMax (564, before the 64-token overlap is prepended) already exceed the window, and a whole-kept code block or table would lose up to three quarters of its content. Silently: an over-long input is truncated, not refused.

So the fixture's cleanliness — longest input 104 tokens, 0 of 60 anywhere near 512 — retires the truncation confound for this fixture and for nothing else. A result measured here licenses no statement about how a 512-token model behaves on production chunk sizes. If anything the chunker constants are an independent and stronger argument against a 512-token embedder than the recall numbers on this page are: the recall gap is a measured tendency, whereas adopting a 512-token model would truncate real content by construction, on the default chunking path, with nothing in the logs to say so.

What the numbers support: "on a fixed 40-document fixture, the shipped local 768-d model placed the intended answer in the top 5 for 20 of 20 paraphrased queries, against 18 of 20 for a same-class 768-d self-hosted alternative and 7 of 20 for a bag-of-words control." Not an absolute quality claim, not an SLA, not a tier or pricing statement, not a comparison against hosted, and not a general statement that one self-hosted model is better than the other — see the two-query caveat above.

Binding 4 — the hosted row: waived, and what stays forbidden​

Recorded as "not measured" with its reason — never blank, never a dash. No comparative claim between local and hosted may appear anywhere, explicit or implied: an unmeasured row plus a comparative adjective is a documented claim with nothing behind it.

G5 was amended and the hosted arm was waived — both ratified by the founder on 2026-09-03T06:30:40Z (5521553923, PRD #996 v2.2). This section previously gated T15 on the hosted arm; that gating language is retired and is recorded here rather than deleted, because the earlier revision of this file is cited elsewhere.

  • G5 now reads as a cross-model comparison across self-hostable models, not hosted-against-local. The delivered local-nomic vs local-arctic evidence satisfies it as amended. T15 — the local-default flip — shipped (#2737, d6eecea4), and it cleared the gate that governed it.
  • The reasoning the founder accepted: T15 changes unset ⇒ refuse into unset ⇒ local. The counterfactual it displaces is refusal, not hosted — nobody is migrated off a hosted embedder because nobody was on one. That is why a hosted comparison was not load-bearing for this flip.
  • The hosted row is WAIVED, not pending. #2795 is closed as waived. The founder declined it as unprovisionable rather than merely unprovisioned: no hosted embedding account exists and LiteLLM holds no upstream credential, so the real cost is a new vendor relationship and a new secret surface, not the ~60 calls. Filing it as "~60 calls, a fraction of a cent" understated the ask, and the PM recorded that correction against themselves.
  • The trigger, for whoever finds this later: if a hosted embedding credential comes to exist for any other reason (Model Gateway or BYOK being the likely source), run the arm — one --arm, one --dim 1536, the provider's own environment. The fixture is now known to discriminate (it separated nomic from arctic at p = 0.0039), so the measurement is known-informative. Record it against G5.

What the waiver bought, and what it permanently costs. The bound below is not a placeholder awaiting the measurement — it was accepted in exchange for the waiver and is therefore permanent unless the trigger above fires:

No comparative claim between local and hosted may appear anywhere, explicit or implied. Say "we did not measure it, and it was measurable" — never "there was nothing to find."

The local-arctic arm does not discharge the hosted question. It is a second self-hosted model, not a hosted one: same endpoint class (e=private), same arity, no provider credential involved. It satisfies G5 as amended and it answers "is there a better self-hosted option we can ship" — a real question, worth 60 calls. It leaves the hosted row unmeasured, and that row is now closed as waived rather than left open.

The control arm, and why the first row means anything​

dev-lexical is the negative control, not a candidate: a working retrieval mechanism — well above the 0.125 chance level, with a distinct top hit for most queries — that is simply blind to paraphrase. It breaches all four floors. Its job is to make a passing verdict falsifiable.

The same separation is asserted in CI, offline and with no database, by internal/context/recallbench.TestLexicalArmFailsTheGateOnTheParaphraseQuerySet. It reddens if the query set ever becomes lexically easy — measured: with every query replaced by its own answer's text, the lexical arm scores 0.800 / 1.000 / 0.927 and the test fails, naming the cause.

Independently reproduced inside the air-gapped stack​

scripts/embedding-local-egress-smoke.sh assertion G runs this benchmark against the same artifact from a network with no default route (Docker internal: true, with a control that dials a public address and requires failure), through retrieval.VectorSearcher — the ANN alone, no BM25, no RRF. That is the second half of #2546's claim: nothing leaves the perimeter, and retrieval still works.

Effect of the adjudication, disclosed​

The PM made the q17/d36 call from the corpus text before opening any result artifact, and disclosed that it moves the headline up. Measured effect:

beforeafterdelta
graded recall@50.8750000.883333+0.008333
MRR@50.8833330.8833330
nDCG@50.8768310.879468+0.002636

Control arm unchanged (0.375 / 0.304 / 0.297, 12/20 zero-recall, still fails all four conditions), so arm separation is preserved.

One divergence from the PM's predicted table, and it is arithmetic, not substance. The ruling predicted nDCG 0.880; the run measured 0.879468. The estimate rounded twice — 0.877 + 0.003 — where the unrounded values are 0.876831 + 0.002636 = 0.879468. Recall and MRR match the prediction exactly. q17's nDCG was also recomputed by hand against the retrieved ranking (d36 at rank 2) and agrees with the engine to six decimal places: 0.878962.

Reproducing​

scripts/embedding-recall-bench.sh \
--arm local-nomic --dim 768 \
--env EMBEDDING_PROVIDER=openai \
--env EMBEDDING_BASE_URL=http://127.0.0.1:11434/v1 \
--env-file deployments/embedder/nomic-embed-text.pin.env

# the challenger — same corpus, same queries, same labels, its own scratch DB
scripts/embedding-recall-bench.sh \
--arm local-arctic --dim 768 \
--env EMBEDDING_PROVIDER=openai \
--env EMBEDDING_BASE_URL=http://127.0.0.1:11434/v1 \
--env-file deployments/embedder/snowflake-arctic-embed-110m.pin.env

# the control — expected to FAIL the gate, and the script requires that it does
scripts/embedding-recall-bench.sh \
--arm dev-lexical --dim 1536 --expect-fail \
--env EMBEDDING_PROVIDER=dev --env ENVIRONMENT=development

Adding the third arm needed no change to the harness — an arm is a pin file plus two settings. Provenance for the challenger is deployments/embedder/snowflake-arctic-embed-110m.pin.env, derived from the registry with scripts/embedding-artifact-pin.sh --show and verified against the live Ollama store with --verify-store before the arm was scored, so the number above names bytes rather than a tag.

Each arm builds its own scratch database at its own arity and re-embeds the corpus with cmd/tools/rag-reembed — the shipped FR-8 tool, not a private embedding path. Re-runnability is what makes ADR-0031 D5.3's refresh cadence affordable: a weights bump changes r=, which invalidates the corpus, and these numbers have to be taken again.