retrieval_min_score — the assembly-leg relevance floor
Operator reference for the Context Engine knob added by core#2618.
Wire contract: docs/api/contextengine.md (generated). This page is the why.
What it does
ContextEngineKnobsConfig.retrieval_min_score is the relevance floor that
ContextEngine.AssembleContext's retrieval leg applies before retrieved chunks
enter the assembled prompt.
| Scale | normalised RRF fused score, (0, 1] |
| Range accepted | [0.0, 1.0] (validated at the RPC; CHECK in migration 226) |
0.0 means | platform default, not "no floor" — same as every sibling knob |
| Platform default | 0.3 (knobsconfig.DefaultRetrievalMinScore) |
| Clearance | edit L4 · read L3 |
| Applies to | AssembleContext only — see Scope |
| Effect latency | next call on the pod (in-process Reader.Put on Update; no restart) |
To set an effectively zero floor, use a small positive value (0.0001, which
is what the retrieval integration tests already do). A literal 0 is the
"defer to the platform" sentinel and cannot mean "no floor" — proto3 gives a
bare double no field presence, and every column on context_engine_knobs
already reads 0 that way.
Why the default is 0.3, and why it is not 0.7
Two facts fix it. Neither is a preference.
1. Arithmetic — a floor above 1/3 excludes vector-only matches outright
internal/context/retrieval/rrf.go normalises the fused score by the summed
weight of the rankers that returned rows:
score(d) = Σ_r w_r/(k + rank_r(d)) ÷ Σ_active w_r/(k + 1)
With the shipped DefaultRRFConfig (equal weights, k = 60) and all three
rankers live, a chunk found by only m of the 3 rankers has a hard ceiling of
m/3:
| found by | ceiling | cleared 0.7? |
|---|---|---|
| vector only (paraphrase match, no lexical overlap) | 0.333 | never |
| vector + BM25 (relevant but not among the most recent) | 0.667 | never |
| all three | 1.000 | only if ranked high in each |
So 0.7 was never a relevance bar. It was a structural requirement that the
chunk also appear in the recency ranker — and RecencySearcher.Search takes
no query and no vector at all; it returns the most recent chunks in scope. The
old default therefore made "is this recent" a precondition for "is this
relevant".
This is the mechanism behind the field report on #2618: "workflow approval gates" peaked at 0.66 — the 2-of-3 band, just under 0.667 — and returned
nothing, while "workflow pause endpoint" scored 0.97/0.91/0.91 (3-of-3) and
worked. A spot-check passes while a large share of real queries silently
degrade to no-context.
Any default > 1/3 makes a vector-only match unreachable no matter how relevant it is — and vector-only paraphrase matches are precisely the ones BM25 misses by construction.
2. Internal consistency — it is the engine's own expansion floor
retrieval.ConfidenceGate already declares MinConfident = 0.5 ("accept
without expanding") and ExpandedFloor = 0.3 ("widen the net and keep what
comes back"). Before this change the final filterByMinScore(results, 0.7) ran
after the expansion pass had already filtered at ExpandedFloor, so
everything expansion admitted was immediately discarded — ExpandedFloor was
dead code on the assembly path.
Setting the assembly floor to ExpandedFloor is the value that makes the
pipeline agree with itself, and 0.3 < 1/3 satisfies the arithmetic bound
above. TestPlatformRetrievalFloorEqualsTheExpansionFloor reddens if either
side moves alone.
What did not set this number
The in-repo recall benchmark (docs/measurements/recall-bench/) cannot set
it, and no figure here is derived from it:
- Binding 3 item 2: scoring there is vector-only, while users are served by
the hybrid stack (ANN + BM25 + RRF + recency). It therefore says nothing about
the fused score distribution that
min_scorefilters. - Binding 2: "We have no product bar for retrieval quality, and this fixture cannot set one."
The benchmark is used here for exactly one claim, which it does support: vector-only retrieval is the mode the shipped local model is good at (primary recall@5 = 1.000 on 20/20 paraphrase queries) — which is what makes excluding vector-only matches a real loss rather than a theoretical one.
core#1410 — "tune min_score/gate defaults against a real recall benchmark" —
therefore remains OPEN. Closing it needs a hybrid-scored fixture that does
not exist yet. This change makes the floor explicit, configurable and
arithmetically defensible; it does not claim to be tuned.
Scope: what this knob does not cover
It governs the assembly leg only. Direct ContextEngine.Search callers are
unchanged and keep the historical behaviour:
SearchRequest.min_score <= 0⇒ the server applies 0.7 (internal/context/retrieval/service.go:282-284,:606-608).
Because proto3 gives min_score no field presence, 0 and unset are the same
byte on the wire, and both mean the strictest floor the engine has — not "no
floor". Adding presence (optional double) would flip the generated Go type
from float64 to *float64 and break every caller across two repos, so it was
deliberately not done here.
Known consequences, stated rather than left to be rediscovered:
- The gwpoc gateway's
-min-score 0.15pin stays required. Seedocs/architecture/gwpoc-inspection-gateway.md§8. #2618 did not retire it, and neither did core#3079. - The Retrieval Test UI still sends no
min_score. The request body QA captured on #1411 is{"scope":…,"query":…,"topK":8}— no floor — so that page inherits 0.7 and remains exposed to this trap for queries whose fused scores land below it. Tracked on #2618; not a #1411 regression. What core#3079 changed for it: the floor is now on the response (effective_min_score), together withresults_below_thresholdandoutcome, so the page can display the floor it silently inherited and tell an empty corpus from a filtered one. The page still has to be changed to read them — this makes the symptom self-describing, it does not fix the request.
Anything that wants the knob's floor should either call AssembleContext, or
pass an explicit min_score of its own.
What the response tells you (core#3079)
Both retrieval surfaces now state what the floor did, so "is this empty because the corpus has nothing, or because the bar is too high" is a field rather than an investigation.
| Field | SearchResponse | QualityAttestation (AssembleContext) |
|---|---|---|
| the floor that ran | effective_min_score | effective_min_score |
| candidates the floor removed | results_below_threshold | results_below_threshold |
| the classification | outcome | retrieval_outcome |
outcome is the closed RetrievalOutcome enumeration. Phase 1 produces three
of its members — ok, no_results, below_threshold; the four F3/F4
identities (access_revoked_upstream, source_unavailable,
document_no_longer_exists, document_ref_unresolved) are declared and not yet
emitted. unspecified means the responder did not classify the call, and on the
assembly surface it means the retrieval leg did not run — fallback_used and
fallback_reason name the cause in every one of those cases.
Two consequences worth knowing before reading the numbers:
results_below_thresholdcounts what the FLOOR removed, out of the fused candidate set retrieval produced. Rows omitted becausetop_kwas smaller than the number available are above the floor and are not counted.- It is identical on a cache hit and a cache miss. The result cache stores the artefact pre-floor (#3000 / core#3086), so both branches count against the same set. A count that differed between the two runs of one query would be worse than no count — it would read as the corpus having changed.
Operating it
# Read (L3)
POST /upsquad.contextengine.v1.ContextEngineKnobsConfigService/GetContextEngineKnobsConfig
{}
# Write (L4) — the payload REPLACES the tenant's knob set; send every field
# you want to keep. Zero-valued fields fall back to platform defaults.
POST /upsquad.contextengine.v1.ContextEngineKnobsConfigService/UpdateContextEngineKnobsConfig
{"config": {"retrievalMinScore": 0.15, "retrievalVectorTopK": 32}}
A successful Update emits one config.updated audit event whose
changed_field_paths includes retrieval_min_score
(TestDiffSnapshotsReportsRetrievalMinScore pins that).
Diagnosing a suspected floor problem
Since core#3079 the response answers this directly — read it before running
anything. AssembleContext's quality block carries all three fields:
retrieval_outcome: below_thresholdwithresults_below_threshold: 12→ the corpus has 12 candidates and the floor removed them. Lower the knob.effective_min_scoreis the floor that did it.retrieval_outcome: no_results→ not a floor problem. Check the fan-out warningretrieval: fan-out produced zero rows or had searcher failures (#2047), which reports per-searcher row counts, per-searcher errors and the effective-source ACL size.retrieval_outcome: unspecifiedwithfallback_used: true→ retrieval did not run.fallback_reasonsays which of the three reasons it was.
The old procedure — reissue the query at a low floor and compare — is kept below, because it is still the way to see the actual score distribution (the counts tell you how many fell under the bar, not by how much):
grpcurl -d '{"query":"<the task text>","topK":10,"minScore":0.01}' \
… upsquad.context.v1.ContextEngine/Search
Note that it is no longer needed to answer the yes/no question, and that it was never sound as a substitute for the counts: a second call runs a second retrieval, so a corpus that changed in between makes the comparison lie.
References
- Issue: core#2618 · contract half: core#3079 (LLD #3072 T7, PRD #2995 TK-5.5) · follow-up: core#1410 · related: core#1401 (Defect C, the RRF normalisation this floor is expressed on top of), core#3000 / core#3086 (the pre-floor cache write that makes the below-threshold count free on a cache hit)
- Code:
internal/context/assembly/pipeline.go(retrievalMinScore),internal/contextengine/knobsconfig/reader.go(DefaultRetrievalMinScore),internal/context/retrieval/rrf.go(Fusenormalisation),internal/context/retrieval/outcome.go(applyFloor— the one place the floor is applied, and where the accounting comes from) - Schema:
internal/context/store/migrations/226_context_engine_knobs_retrieval_min_score.up.sql