ADR-0036 — A non-streaming caller gets its own upstream head wait
- Status: Accepted (implementation record). The decision is the architect's ruling on #3620 (comment 5839683504) — "Option B now, in modified form; Option A (streaming) is explicitly NOT scheduled" — with its three conditions. This ADR records what was decided and, more importantly, the numbers and the two refusals, so neither has to be re-derived from a PR diff.
- Date: 2026-09-26
- Deciders: Principal Architect (ruling on #3620, three conditions). Backend SME (the number, and the two bounds on the override). No founder sign-off required — no auth primitive, no secret surface, no egress change, no dependency, no migration.
- Refs: #3625 (this work) · #3620 (the report and the ruling) · #3623 (the diagnostics half, which this changes the bound of) · #2819 / case C34 (the invariant this re-uses) · #2813 (
MaxRequestDurationis derived and has no knob) · ADR-0034 (the Model Gateway) · LLD #2801 §8 - Numbering note: ADR-0032 through ADR-0035 are claimed by issues (#2599, #2601, #2643, #3063) and have no file in
docs/adrs/. 0036 is the next unclaimed number; the gap is pre-existing and is not filled here.
Context
internal/modelgateway/build.go built ONE outbound client and set
tr.ResponseHeaderTimeout = upstreamDialTimeout (30s) on it. The constant's own
doc argued the case correctly — and only for a streamer:
upstreamDialTimeoutbounds the wait for an upstream RESPONSE HEAD only. It is emphatically NOT a whole-request timeout: an LLM turn streams for minutes… So the client's Timeout stays ZERO and the bound is applied where it belongs, to the connection and the response header.
For stream: true the head arrives in milliseconds and thirty seconds is
generous. For stream: false the head IS the answer — nothing reaches the
wire until the completion is finished — so the same thirty seconds was a hard
generation budget the caller never asked for, and the cut surfaced through
handler.go's pre-response branch as upstream_unavailable / 502: a claim
about the TENANT'S provider, made about a cut this gateway performed.
Measured on beta-dev 2026-09-25, upsquad-ai/upsquad-core (6,320 manifest
entries, ~76k prompt tokens), same model and endpoint four minutes apart:
| run | upstream | total_ms | outcome |
|---|---|---|---|
| 19:03:59 | 200 | 25,833 | under the wire |
| 19:07:33 | — | 30,094 | cut |
A 151-entry repository on the same model takes 18.3s and succeeds. The largest repository registered sat on the boundary, so publishing its map was a coin flip.
This is the repo's dominant defect class, again: a bound chosen for one
purpose silently governing another. upstreamDialTimeout governs two phases
(the TCP dial and the response head) and was reasoned about for one caller.
Decision
D1 — Two head waits, selected by request mode
transport.HTTPTransport holds a Clients pair. transport.Request carries
StreamRequested, decided ONCE by the resolver from the one body read, carried
on passthrough.Route, and copied into the transport request. The transport
selects the client; nothing else about the request changes.
The dial bound did not move and must not. Thirty seconds to complete a TCP handshake is already an eternity and nothing about a non-streaming body makes a connection slower to establish. Only the response-head half is mode-dependent.
The mode is the RESOLVER's determination, not the handler's. Under
ModeTranslate the body the handler holds at the dial has been REPLACED with
bytes this gateway wrote; a handler that re-derived the mode there would be
deciding how long to wait for a tenant on the strength of its own output.
false is the safe default and it is also the zero value. A caller that
never sets the field gets the PATIENT bound. A streamer wrongly given the
patient bound waits longer for a head that arrives in milliseconds anyway and is
still cut by the handler's whole-request cap; a non-streamer wrongly given the
strict bound is this defect, back again. Only one of those two directions can
be re-created by omission, and it is the harmless one.
D2 — The number is 90 seconds, and it is NOT derived from the drain
This is the architect's first condition and it is the substantive one.
MaxRequestDuration is derived from DrainTimeout because it is a
drain-safety bound — a statement about the pod's lifecycle. This is not: it is a
statement about how long a model legitimately takes to answer. Deriving one from
the other would be the same error as the defect.
Ninety seconds, for three reasons and no fourth:
- Three times the streaming head wait. Two bounds differing by a factor of three cannot be confused by a reader, and a mutation swapping one for the other changes an observable.
- ~3.5× the largest non-streaming completion this platform has been measured producing (25,833 ms, above). A provider three and a half times slower than the slowest we have actually observed still completes.
- It bounds ONE completion, not a session. A provider that cannot answer a single chat completion inside a minute and a half is not about to.
Fit against the fleet, checked AFTER the choice and never used as its source: the tightest derived cap is the dev overlay's 150s (drain 180 − margin 30), whose allowance is 150 − 30 = 120s. Ninety clears it with a whole margin to spare — and the boundary is precisely where #2819's C34 lived. Had every drain in the fleet been 900s, the number would still be 90.
What it does not buy, stated rather than implied: 8000 output tokens in 90 seconds is a sustained ~89 tok/s, which a loaded self-hosted server will not reach. That is what the override is for, which is why it is documented as an override and not presented as a tuning dial.
D3 — Two refusals, both at startup, both reachable
MODEL_GATEWAY_UPSTREAM_NON_STREAMING_HEAD_TIMEOUT_SECONDS overrides the
default. LoadConfig refuses a configuration where the effective value:
- is not positive — it would cut every non-streaming call instantly from a gateway whose startup log reports the bound as configured;
- is below
UpstreamHeadTimeout— a non-streamer is strictly MORE patient than a streamer by construction, so a shorter bound gives the caller that needs the most time the least. That is this defect re-entered through a configuration file, and there is no deployment for which it is right; - does not clear
MaxRequestDurationby a wholeMaxRequestMarginSeconds— the head wait is spent INSIDE the cap (handler.gowraps the upstream call in it deliberately), so at less clearance the CAP cuts the head wait andfailreportsupstream_unavailable. That is #2819 case C34 exactly, with the other head wait in it.
A FULL margin of clearance rather than one second, and the extra is not superstition: the remainder is what the response body and the pipe run in after the head arrives.
Consequence, stated because it is a real tightening: the smallest drain that
loads at all rises from 31s to 150s (90 + 30 + 30). Every manifest in this
repo is 180 or 900, so no deployment moves; a hand-built Config in a test with
a smaller cap must set the field itself (model_gateway_pool_exhaustion_3061
does, to 1s, and says why).
D4 — The CI guard reads BOTH constants, and C34 had to be re-armed
scripts/check-model-gateway-exposure.py read ONE head-wait constant by literal
name. A non-streaming bound at or above the cap reproduces C34 while that
check stays green, because its regex does not know the new constant exists.
The guard now reads DefaultUpstreamNonStreamingHeadTimeout and any per-variant
manifest override, and asserts the relation on the EFFECTIVE value per variant.
Cases C119–C124 sit beside C34/C35.
Two consequences of the new floor that had to be handled rather than papered
over, because the new relation SUBSUMES the old one arithmetically (ns_head ≥ head_wait and margin > 0 ⟹ nothing can clear the new floor and fail the old):
- C35's original premise is no longer a passing tree. Drain 61 (cap 31) really does clear the 30s streaming floor by one second, and is then refused by the second floor. It now asserts exactly that — rc=1, naming the non-streaming bound and NOT the streaming one — which is the same anti-vacuity property restated for the world it now lives in.
- C34 now greps for the streaming arm's own message. Exit 1 alone no longer proves which arm fired, so without the grep the streaming floor could be deleted outright with nothing red — a check made unfalsifiable by its successor.
The same treatment was needed in config_test.go
(drain=31 is the smallest that is accepted → drain=31 clears THIS arm and is refused by the head-wait one, plus a new drain=149/150 pair).
D4.1 — The bound set is DECLARED ONCE, and the guard reads it in reverse
Reading both constants by name fixed those two constants and left the hole open
for the next one. The architect proved it on review: a third member on
transport.Clients — Probe *http.Client built at 600 s, four times the dev
cap and #2819 case C34 exactly for that member — and the checker returned 0,
both Go packages passed and the mutation battery ran 127 green. Hand-
enumeration was the defect, not the missing name.
The set is declared in exactly one place: modelgateway.upstreamHeadBounds
(internal/modelgateway/build.go) — one entry per transport.Clients member,
naming the declaration its duration comes from. Nothing else enumerates the
members: assembleUpstreamClients walks the struct and refuses to build a
member with no entry, and transport.eachMember / transport.ClientsForEveryMode
replaced the other four hand-written lists (the default pair, the nil backfill,
the redirect policy, the probe's same-client-twice).
The register is then checked in reverse by two independent oracles — the Go
suite over reflect, and check_upstream_head_bounds_are_registered in the
exposure script:
| Tree | Verdict |
|---|---|
| member added, not registered | startup refusal + exposure checker exit 1, naming the member |
| registered, member gone | startup refusal + exit 1 on the stale entry |
| registered against a source the script cannot value | exit 2 — "refusing to guess" |
| register entry whose resolver disagrees with its own source | Go suite red |
| head wait assigned anywhere else in the package | exit 1, unless it is a declared exclusion |
transport.defaultResponseHeaderTimeout is the one named exclusion, recorded
with its reason in UNREGISTERED_HEAD_WAIT_EXCLUSIONS and on the constant: it
is the 60 s fallback for a nil member, and it is unreachable from the gateway
because every member is filled from the register before transport.New sees the
pair. The exclusion is "unreachable on the gateway's path", not "small enough
not to matter".
Cases C125–C133 are the kills, C125 being the architect's own mutant.
D5 — repomapgen reads the bound that ACTUALLY APPLIED to it
classifyGatewayNon200 split "the provider timed out" from "the provider was
never reached" on modelgateway.UpstreamHeadTimeout. repomapgen sends
"stream": false, so the moment D1 lands that is the wrong constant: every head
timeout between 30s and 90s would read as provider_unreached — a refused dial,
which sends an operator to look at egress rules for a problem that is a slow
model. That is #3620's own thesis failing inside the code written to implement
it, which is why it is a condition of this change rather than a follow-up.
It now reads DefaultUpstreamNonStreamingHeadTimeout.
The residual, stated rather than hidden. That is the gateway's DEFAULT, and
a deployment may override it. Raising the override leaves the classifier sound
(a head timeout at a raised bound still exceeds the default, so it still lands
on provider_timeout); lowering it makes the split approximate in one
direction. No manifest in this repo sets the variable. The honest fix is the
gateway recording the bound it applied on the ledger row the classifier already
fetches — that is a migration, and it is deliberately outside #3625.
Rejected
Streaming (Option A) — NOT SCHEDULED, and deliberately not "later"
The architect's ruling, recorded here so it is not re-litigated: "B now, A later" schedules A by default, and after B its remaining value to repo-map is close to nil — a background worker, nobody watching tokens, no cancellation story. A stays open as a gateway/client capability whose first consumer should be an interactive surface.
It must also not be the vehicle that carries the include_usage problem into
production: dialect/openai.go names vLLM, Together, Groq, Fireworks, Azure
OpenAI, LiteLLM and Ollama as servers that ignore the flag, and on those,
streaming converts "sometimes times out" into "never publishes".
Central stream_options.include_usage injection — REFUSED as a body mutation
The bytes that were governed must be the bytes that were sent. dialect/openai.go
claimed the transport already asks for usage; it does not, and nothing in this
tree sends the flag. The sentence is corrected here; the behaviour is not
changed to make a wrong sentence right. If the requirement is real it belongs at
admission, as its own issue with its own ADR.
A per-request context deadline instead of a second client
ResponseHeaderTimeout is a field on http.Transport, not a per-request value.
The only per-request alternative is a context deadline, which covers the
response BODY as well as the head — on a streaming request that is precisely the
guillotine the gateway's Timeout: 0 rule exists to avoid, and the handler
already owns the one whole-request deadline there should be.
Clamping the bound to MaxRequestDuration − margin instead of refusing
Silently substituting a value an operator did not choose, and — worse — it would make the bound derived from the drain window, which is exactly what D2 forbids.
Consequences
- A non-streaming completion may now take up to 90 seconds (default) instead of 30. The whole-request cap is unchanged and still binds first on any variant whose cap is below the bound.
- The dev overlay's 150s cap may still bind.
modelStageasks for 8000 output tokens; below ~67 tok/s end-to-end the cap, not the head wait, is the limit. The measured failures were head timeouts at exactly 30s, so this should suffice — but it is re-measured on beta-dev and recorded on #3625 rather than assumed. If dev still binds, the answer is one configmap line raising the drain, not Option A. - Two
http.Transports means two connection pools (up to 32 idle conns per host each instead of 32 total). A cost in sockets, not a correctness question. - Both clients carry the identical dial-time egress fence. A second client is
a second place for
tr.DialContext = guard.DialContextto go missing, and the missing-guard state is indistinguishable from the guarded one until somebody registershttp://169.254.169.254/.TestUpstreamClients_TheEgressFenceIsOnBOTHModesClientsdrives a real socket and a real refusal through each member. - No migration. No change to the ledger schema, the refusal vocabulary, or any wire contract.