Skip to main content

retrieval_min_score — the assembly-leg relevance floor

Operator reference for the Context Engine knob added by core#2618. Wire contract: docs/api/contextengine.md (generated). This page is the why.

What it does​

ContextEngineKnobsConfig.retrieval_min_score is the relevance floor that ContextEngine.AssembleContext's retrieval leg applies before retrieved chunks enter the assembled prompt.

Scalenormalised RRF fused score, (0, 1]
Range accepted[0.0, 1.0] (validated at the RPC; CHECK in migration 226)
0.0 meansplatform default, not "no floor" — same as every sibling knob
Platform default0.3 (knobsconfig.DefaultRetrievalMinScore)
Clearanceedit L4 · read L3
Applies toAssembleContext only — see Scope
Effect latencynext call on the pod (in-process Reader.Put on Update; no restart)

To set an effectively zero floor, use a small positive value (0.0001, which is what the retrieval integration tests already do). A literal 0 is the "defer to the platform" sentinel and cannot mean "no floor" — proto3 gives a bare double no field presence, and every column on context_engine_knobs already reads 0 that way.

Why the default is 0.3, and why it is not 0.7​

Two facts fix it. Neither is a preference.

1. Arithmetic — a floor above 1/3 excludes vector-only matches outright​

internal/context/retrieval/rrf.go normalises the fused score by the summed weight of the rankers that returned rows:

score(d) = Σ_r w_r/(k + rank_r(d)) ÷ Σ_active w_r/(k + 1)

With the shipped DefaultRRFConfig (equal weights, k = 60) and all three rankers live, a chunk found by only m of the 3 rankers has a hard ceiling of m/3:

found byceilingcleared 0.7?
vector only (paraphrase match, no lexical overlap)0.333never
vector + BM25 (relevant but not among the most recent)0.667never
all three1.000only if ranked high in each

So 0.7 was never a relevance bar. It was a structural requirement that the chunk also appear in the recency ranker — and RecencySearcher.Search takes no query and no vector at all; it returns the most recent chunks in scope. The old default therefore made "is this recent" a precondition for "is this relevant".

This is the mechanism behind the field report on #2618: "workflow approval gates" peaked at 0.66 — the 2-of-3 band, just under 0.667 — and returned nothing, while "workflow pause endpoint" scored 0.97/0.91/0.91 (3-of-3) and worked. A spot-check passes while a large share of real queries silently degrade to no-context.

Any default > 1/3 makes a vector-only match unreachable no matter how relevant it is — and vector-only paraphrase matches are precisely the ones BM25 misses by construction.

2. Internal consistency — it is the engine's own expansion floor​

retrieval.ConfidenceGate already declares MinConfident = 0.5 ("accept without expanding") and ExpandedFloor = 0.3 ("widen the net and keep what comes back"). Before this change the final filterByMinScore(results, 0.7) ran after the expansion pass had already filtered at ExpandedFloor, so everything expansion admitted was immediately discarded — ExpandedFloor was dead code on the assembly path.

Setting the assembly floor to ExpandedFloor is the value that makes the pipeline agree with itself, and 0.3 < 1/3 satisfies the arithmetic bound above. TestPlatformRetrievalFloorEqualsTheExpansionFloor reddens if either side moves alone.

What did not set this number​

The in-repo recall benchmark (docs/measurements/recall-bench/) cannot set it, and no figure here is derived from it:

  • Binding 3 item 2: scoring there is vector-only, while users are served by the hybrid stack (ANN + BM25 + RRF + recency). It therefore says nothing about the fused score distribution that min_score filters.
  • Binding 2: "We have no product bar for retrieval quality, and this fixture cannot set one."

The benchmark is used here for exactly one claim, which it does support: vector-only retrieval is the mode the shipped local model is good at (primary recall@5 = 1.000 on 20/20 paraphrase queries) — which is what makes excluding vector-only matches a real loss rather than a theoretical one.

core#1410 — "tune min_score/gate defaults against a real recall benchmark" — therefore remains OPEN. Closing it needs a hybrid-scored fixture that does not exist yet. This change makes the floor explicit, configurable and arithmetically defensible; it does not claim to be tuned.

Scope: what this knob does not cover​

It governs the assembly leg only. Direct ContextEngine.Search callers are unchanged and keep the historical behaviour:

SearchRequest.min_score <= 0 ⇒ the server applies 0.7 (internal/context/retrieval/service.go:282-284, :606-608).

Because proto3 gives min_score no field presence, 0 and unset are the same byte on the wire, and both mean the strictest floor the engine has — not "no floor". Adding presence (optional double) would flip the generated Go type from float64 to *float64 and break every caller across two repos, so it was deliberately not done here.

Known consequences, stated rather than left to be rediscovered:

  • The gwpoc gateway's -min-score 0.15 pin stays required. See docs/architecture/gwpoc-inspection-gateway.md §8. #2618 did not retire it, and neither did core#3079.
  • The Retrieval Test UI still sends no min_score. The request body QA captured on #1411 is {"scope":…,"query":…,"topK":8} — no floor — so that page inherits 0.7 and remains exposed to this trap for queries whose fused scores land below it. Tracked on #2618; not a #1411 regression. What core#3079 changed for it: the floor is now on the response (effective_min_score), together with results_below_threshold and outcome, so the page can display the floor it silently inherited and tell an empty corpus from a filtered one. The page still has to be changed to read them — this makes the symptom self-describing, it does not fix the request.

Anything that wants the knob's floor should either call AssembleContext, or pass an explicit min_score of its own.

What the response tells you (core#3079)​

Both retrieval surfaces now state what the floor did, so "is this empty because the corpus has nothing, or because the bar is too high" is a field rather than an investigation.

FieldSearchResponseQualityAttestation (AssembleContext)
the floor that raneffective_min_scoreeffective_min_score
candidates the floor removedresults_below_thresholdresults_below_threshold
the classificationoutcomeretrieval_outcome

outcome is the closed RetrievalOutcome enumeration. Phase 1 produces three of its members — ok, no_results, below_threshold; the four F3/F4 identities (access_revoked_upstream, source_unavailable, document_no_longer_exists, document_ref_unresolved) are declared and not yet emitted. unspecified means the responder did not classify the call, and on the assembly surface it means the retrieval leg did not run — fallback_used and fallback_reason name the cause in every one of those cases.

Two consequences worth knowing before reading the numbers:

  • results_below_threshold counts what the FLOOR removed, out of the fused candidate set retrieval produced. Rows omitted because top_k was smaller than the number available are above the floor and are not counted.
  • It is identical on a cache hit and a cache miss. The result cache stores the artefact pre-floor (#3000 / core#3086), so both branches count against the same set. A count that differed between the two runs of one query would be worse than no count — it would read as the corpus having changed.

Operating it​

# Read (L3)
POST /upsquad.contextengine.v1.ContextEngineKnobsConfigService/GetContextEngineKnobsConfig
{}

# Write (L4) — the payload REPLACES the tenant's knob set; send every field
# you want to keep. Zero-valued fields fall back to platform defaults.
POST /upsquad.contextengine.v1.ContextEngineKnobsConfigService/UpdateContextEngineKnobsConfig
{"config": {"retrievalMinScore": 0.15, "retrievalVectorTopK": 32}}

A successful Update emits one config.updated audit event whose changed_field_paths includes retrieval_min_score (TestDiffSnapshotsReportsRetrievalMinScore pins that).

Diagnosing a suspected floor problem​

Since core#3079 the response answers this directly — read it before running anything. AssembleContext's quality block carries all three fields:

  • retrieval_outcome: below_threshold with results_below_threshold: 12 → the corpus has 12 candidates and the floor removed them. Lower the knob. effective_min_score is the floor that did it.
  • retrieval_outcome: no_results → not a floor problem. Check the fan-out warning retrieval: fan-out produced zero rows or had searcher failures (#2047), which reports per-searcher row counts, per-searcher errors and the effective-source ACL size.
  • retrieval_outcome: unspecified with fallback_used: true → retrieval did not run. fallback_reason says which of the three reasons it was.

The old procedure — reissue the query at a low floor and compare — is kept below, because it is still the way to see the actual score distribution (the counts tell you how many fell under the bar, not by how much):

grpcurl -d '{"query":"<the task text>","topK":10,"minScore":0.01}' \
… upsquad.context.v1.ContextEngine/Search

Note that it is no longer needed to answer the yes/no question, and that it was never sound as a substitute for the counts: a second call runs a second retrieval, so a corpus that changed in between makes the comparison lie.

References​

  • Issue: core#2618 · contract half: core#3079 (LLD #3072 T7, PRD #2995 TK-5.5) · follow-up: core#1410 · related: core#1401 (Defect C, the RRF normalisation this floor is expressed on top of), core#3000 / core#3086 (the pre-floor cache write that makes the below-threshold count free on a cache hit)
  • Code: internal/context/assembly/pipeline.go (retrievalMinScore), internal/contextengine/knobsconfig/reader.go (DefaultRetrievalMinScore), internal/context/retrieval/rrf.go (Fuse normalisation), internal/context/retrieval/outcome.go (applyFloor — the one place the floor is applied, and where the accounting comes from)
  • Schema: internal/context/store/migrations/226_context_engine_knobs_retrieval_min_score.up.sql