Connecting an external agent to the Model Gateway
Operator onboarding — MG-3.9. How an agent that runs outside UpsQuad obtains a platform token and completes a governed LLM call through the Model Gateway, what it will be refused for, and where the resulting spend shows up.
Read this before you read anything else
The hostname is LIVE. Production tenant LLM traffic is still gated. Those are two different statements and the difference is the whole of this box.
https://model-gw-beta.upsquad.airesolves, serves and is sealed, verified from the public internet on 2026-09-07 — the battery is in §2 and on the flip comment. What it is released for is E2E testing traffic under the scoped waiver recorded on #2912; production tenant LLM traffic remains gated — #3120 is the named open gate, and the standing list is the query in §9, not a copy on this page. Do not read "the endpoint answers" as "the platform is carrying tenant load" (LLD #2901 §0.1).The hostname took two renames to get here and the reason is worth one line, because it constrains the next one:
model-gateway.app.upsquad.ai→model-gw-beta.app.upsquad.ai(PR #3196) →model-gw-beta.upsquad.ai(PR #3198). A two-label name needs a per-hostname certificate — Cloudflare's universal cert is*.upsquad.ai+upsquad.ai, and a TLS wildcard matches exactly one label — so with Total TLS off the cert never issued. The one-label name was covered instantly. Any future MG hostname must stay one-label underupsquad.aiunless Total TLS is turned on first.So: §2 has a devbox / loopback half and a public hostname half, and both are now live measurements rather than one measurement and one plan. Every control below names the task that delivered it, so you can re-verify it rather than re-trust this page.
1. What you are connecting, and who does what
An external agent is a first-class agents.id — team, clearance, per-agent model
enablement, budget and audit all apply to it identically to a platform agent. The only
difference is the runtime = "external" provenance stamped on its token and carried onto
every ledger and audit row (internal/auth/oauth/token_endpoint.go, MG-3.3 / T5 #2907).
There is no second registration concept for LLM access. The same
external_agent_clients row (migration 155, #1956) that authenticates a BYO agent to the
governed MCP gateway authenticates it to the Model Gateway; only the requested audience
differs (MG-3.3).
Onboarding is five writes, and they are deliberately split across three authorities:
| # | Step | Surface | Who | Delivered by |
|---|---|---|---|---|
| 1 | Register the upstream endpoint | AIGatewayEndpointRegistryService.RegisterLLMEndpoint | L4 | slice 1 T9 #2685 |
| 2 | Activate it | SetLLMEndpointStatus → approved | L5 | slice 1 T9 #2685 |
| 3 | Write the provider credential | AIGatewayEndpointCredentialService.SetLLMEndpointCredential | L4/L5 per grain | slice 1 T10 #2686 |
| 4 | Bind the endpoint to a team, then approve the binding | CreateLLMBinding → ApproveLLMBinding | L4 + the unit's Manager | slice 2 T2 #2807 |
| 5 | Register the external agent + narrow its model access | AgentService.RegisterExternalAgent | L4 | #1956 + slice 2 T4 #2810 |
Why registration and activation are different clearances. base_url is an egress
destination. Authoring one is L4; turning it into a destination the platform will dial on
your behalf is an L5 decision. This split is normative and is not relaxable to
auto-approve (proto/upsquad/aigateway/v1/endpoint.proto).
Why binding approval is L4-and-Manager rather than L5. A binding grants an
already-approved endpoint to one unit, and the authority over a unit is its Manager
(org_units.manager_member_id). A unit with no Manager fails closed — there is no
authority to approve, so nobody may.
1.1 The endpoint URL fence — read this before you pick a base_url
base_url is tenant-supplied, so it goes through the full SSRF fence
(internal/modelgateway/registry/baseurl.go, LLD #2673 §8.3). Four rules, applied in
order; every refusal message begins base_url rule N: so you can tell which one rejected
the value.
| Rule | Requirement |
|---|---|
| 1 | Absolute http/https URL with a host |
| 2 | https only. Plaintext http is permitted solely for hosts in MODEL_GATEWAY_DEV_HTTP_HOSTS, which is empty in production |
| 3 | No userinfo component, no fragment |
| 4 | No private, loopback, link-local, CGNAT or ULA destination — unconditionally. Every A/AAAA record is tested; one blocked record refuses the URL. A host that cannot be resolved is refused, because unproven must not read as allowed |
Rule 4 has no escape hatch. MODEL_GATEWAY_DEV_HTTP_HOSTS relaxes the scheme check
and nothing else; it is currently inert and setting it changes nothing
(docker-compose.dev.yml, LLD §8.3 v1.8). This is the single most common surprise when
standing up a shim — see §7.
Not closed in v1, and stated rather than assumed away: DNS rebinding between validation and dial. A hostname that resolves public here can resolve to loopback when the gateway actually dials it. Closing it needs the pinned-resolver work the MCP egress ADR owns (LLD #2673 §13).
2. Reaching the gateway
Today — beta-dev / the devbox stack · verified 2026-09-07
The gateway process runs and serves, but only on loopback. There is no edge route.
| What | Where | Verified |
|---|---|---|
| Model Gateway | http://127.0.0.1:8091 on the devbox — loopback only, no host publish | yes |
| Token endpoint | http://127.0.0.1:8083/oauth/token (the context-engine container's public mux) | yes |
| AS metadata | http://127.0.0.1:8083/.well-known/oauth-authorization-server | yes |
| JWKS | http://127.0.0.1:8083/.well-known/jwks.json | yes |
| LLM audience in force | http://context-engine:8080/llm | yes — read off the process |
Get there over SSH to the devbox (docs/runbooks/devbox-ssh-warp.md). beta-dev.app.upsquad.ai
is a different vhost — the portal — and it sits behind Cloudflare Access; an unauthenticated
request to it returns 302, not the gateway.
The process states its whole posture in one startup line. This is the artefact to read before debugging anything, and it is the source of the numbers on this page:
$ docker logs upsquad-model-gateway | grep -m1 'controls wired'
… gates="7a/clearance_floor -> 7b/agent_enablement -> 7c/model_allowlist -> 7d/budget -> 7e/guardrail"
rate_limits="org 200/s burst 400 -> team 100/s burst 200 -> agent 20/s burst 40"
dialects_relayable="[anthropic openai_compatible]" credentials_enabled=false
org_token_cap_enforced=true production_routing_released=true
agent_token_auth=true agent_token_audience=http://context-engine:8080/llm
agent_scoped_keys=refused environment=development
Read it off the process you are talking to, not off this page. The excerpt above was
taken from the devbox container on 2026-09-07 and the fields quoted are stable across
builds, but the controls= and seam_types= lists grow with each task and the devbox
image lags main. A build's own line is authoritative for that build; this one is an
example of where to look.
Two lines of that are load-bearing for anyone testing here:
credentials_enabled=false— the devbox has no credential store bound, so a call that passes every gate is refused at step 8 withcredential_store_unbound(503). That is the stack, not your configuration.agent_scoped_keys=refused— theuq_key_*deprecation window is closed on this stack (T17 #2919, merged #3068). An agent-scoped API key is refused withagent_scoped_key_demoted(403). OAuth is the only agent path here.
The public hostname — LIVE, verified 2026-09-07
| What | Where | State |
|---|---|---|
| Model Gateway | https://model-gw-beta.upsquad.ai/v1/… | live, edge installed at 2ec6629e3 (#2912) |
| Sealed by | catch-all direct_response: 404 — only /v1/ is routed | live, probed |
| Released for | E2E testing traffic, scoped waiver on #2912 | production tenant traffic still gated (#3120) |
The vhost routes only /v1/ and seals everything else. That seal is the point:
cmd/model-gateway's mux serves GET /metrics, /readyz and /healthz with no
authentication on the same listener, so a prefix: "/" route would have published all
three. Under #2985's ruling — option (b), gateway-enforced, so there is no
Cloudflare Access application in front — that seal is not defence in depth. It is the only
thing between those three handlers and the internet.
So it was probed rather than asserted. From the public internet, 2026-09-07:
POST /v1/chat/completions (no auth) -> 401 {"error":{"message":"missing or malformed
POST /v1/chat/completions (Bearer nope) -> 401 Authorization header","type":"invalid_api_key"}}
GET /metrics -> 404 ┐
GET /readyz -> 404 │ the direct_response seal
GET /healthz -> 404 │
GET / -> 404 ┘
beta-dev.app.upsquad.ai -> 302 · gw-dev.upsquad.ai -> 302 neighbours unaffected
Two details make that battery mean something, and a bare status code would not.
- The 404s carry no
x-envoy-upstream-service-timeheader; the 401 does. So the 401 is the gateway refusing through Envoy, and the four 404s never reached an upstream at all — they are Envoy'sdirect_response. A status code alone cannot tell "the seal fired" from "the gateway happened to 404". - The same
/metricsis200on loopback (curl 127.0.0.1:8091/metrics). The seal is measurably converting a live 200 into a 404, which is the claim; four 404s on their own are equally consistent with a broken config.
The 401 body is the gateway's own JSON and there is no Access interstitial, no redirect — the
identity decision is being taken by T4 #2906's middleware (cmd/model-gateway/main.go),
which is the mechanism ruling (b) selected, observed working from outside.
Rollback is three independent one-minute levers (any one of them takes the hostname
down): delete the cloudflared ingress block and restart cloudflared-upsquad; delete the
CNAME; restore /etc/envoy/backups/envoy.yaml.<TS>.bak and restart envoy. The operational
checklist lives on #2912 and is not restated here — a restatement is a copy, and
the copy is the half that goes stale.
The trap: the audience is not derived from the gateway's hostname
LLMAudienceURI(publicBaseURL) = publicBaseURL + "/llm", wherepublicBaseURLis the platform base URL (MCP_GATEWAY_PUBLIC_BASE_URL) — the same sourceTeamAudienceURIandPlatformAudienceURIuse, and distinct from both so an MCP token can never satisfy the Model Gateway (LLD §3.5, T4 #2906).Requesting
resource=https://model-gw-beta.upsquad.ai/llmdoes not work — not before the flip, not now that the hostname is live, and the two renames it took to get to this name are the argument. Derivingaudfrom the gateway's hostname was considered and refused: the minter would then need to know that hostname, and a hostname change would silently invalidate every outstanding token with a 401 that names nothing. That hypothetical happened twice in one day (#3196, #3198) and cost this contract nothing, which is the design working.Derive it, never guess it: read
issuerfrom the authorization server's metadata and append/llm. That is the same derivation the validator performs.
3. Getting a token · verified 2026-09-07
grant_type=client_credentials only. No refresh tokens — the agent re-grants. TTL is
≈10 minutes. Two client-authentication methods, both advertised in the AS metadata:
private_key_jwt (RFC 7523, preferred — you hold the private key, we store only the public
JWK) and client_secret_post (fallback, bcrypt).
The mechanics of both, the hardening posture (uniform 401 invalid_client for every
failure, per-IP and per-client rate limits, the external_token_issued /
external_token_denied audit rows) and the MCP_GATEWAY_ENABLED / MCP_GATEWAY_TENANTS
enable gate are documented once, in
Platform Token Issuer. Everything there applies unchanged.
What follows is the one thing that differs for the Model Gateway.
The resource parameter is the whole difference
BASE=http://127.0.0.1:8083 # the PLATFORM origin, not the gateway's
# Derive the LLM audience from the issuer — never hand-write it.
LLM_AUD="$(curl -s "$BASE/.well-known/oauth-authorization-server" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["issuer"].rstrip("/")+"/llm")')"
curl -s -X POST "$BASE/oauth/token" \
-d grant_type=client_credentials \
-d "client_id=$CLIENT_ID" -d "client_secret=$CLIENT_SECRET" \
-d "resource=$LLM_AUD"
# => {"access_token":"<jwt>","token_type":"Bearer","expires_in":600}
RFC 8707 resource selects which audience is minted. It is matched exact, never by
prefix — a prefix match would let …/llm-evil select the LLM audience.
resource | Result | Verified |
|---|---|---|
| absent | MCP-gateway audience (the pre-#2906 shape; every existing client is unaffected) | yes — 401 invalid_client for a bogus client, i.e. the resource was accepted |
| the exact LLM audience | Model Gateway audience | yes — 401 invalid_client, i.e. accepted |
| anything else | 400 invalid_target — "the requested resource is not served by this authorization server" | yes |
An unrecognised resource is refused rather than ignored. Ignoring it would hand back an MCP-audience token to a client that asked for the LLM one, which then fails one hop away as a 401 that can only say "wrong audience".
What was NOT executed here. The devbox has zero rows in
external_agent_clientsand zero inllm_endpoints, so no successful mint and no successful relay were performed. Every refusal shape above is a real response from the running stack; the success shapes are read from the source and the Platform Token Issuer runbook, and are marked as such.
4. Making the call
# PUBLIC (live, verified 2026-09-07): MG=https://model-gw-beta.upsquad.ai
# DEVBOX LOOPBACK (unsealed — serves /metrics, /readyz, /healthz):
# MG=http://127.0.0.1:8091
curl -s -X POST "$MG/v1/chat/completions" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"model":"<a model id on an endpoint your team is bound to>",
"messages":[{"role":"user","content":"hello"}]}'
Things worth knowing before the first call:
- Everything under
/v1/is relayed. The gateway does not enumerate upstream routes — a registeredbase_urlmay carry an Azure-style deployment prefix and your path is joined onto it. So/v1/embeddings,/v1/rerankand provider-native paths reach the upstream; a fixed route table would break the endpoints the registry exists to support. - Your credentials are stripped.
Authorization,x-api-key,Api-Key,Cookieand everyX-Upsquad-*header are removed before forward, and the tenant's credential is injected:Authorization: Bearer …foropenai_compatible,x-api-key: …foranthropic(internal/modelgateway/transport/headers.go). - Dialect is inferred from the request shape, not from a route table (T6 #2908). A body
with
system/top_k/thinkingnamesanthropic; one withfrequency_penalty/response_format/seednamesopenai_compatible. A request naming two is refuseddialect_ambiguous(400). A request naming none is not refused — it falls to pass-through, which is what every earlier build did with it. - Cross-dialect translation is opt-in twice. The team's binding must carry
allow_translation, and the request must sendX-Upsquad-Translate: 1. Never inferred: PRD §11 forbids silent translation and inference is exactly that silence. See §4.1 for what happens when it is on. - The credential grain walk is most-specific-wins with no fallback where the binding
forbids one. If the binding sets
credential_mode = per_agentwithrequire_per_agent, there is deliberately no fallback to the team or org key (MG-2.2); the refusal isagent_credential_required. - HTTP/1.1 only.
:8091does not terminate h2c. An h2-terminating proxy in front is fine.
4.1 Translation: anthropic ⇄ openai_compatible
Both directions, request and response, shipped by T7 #2909 (merged 2026-09-07). Two dialects only — no third, no partial-fidelity mode, and no fallback to pass-through on a translation failure: a failure is a refusal.
The rule that makes it honest: a request feature with no cross-dialect analogue is
REFUSED, never dropped. A dropped tool_choice, thinking block or response_format
produces a successful call with different semantics, which is the worst failure shape
available because you are never told. The lossy set is derived from the dialect structs
rather than hand-listed, so a field a vendor ships tomorrow refuses by default.
| What happened | Wire |
|---|---|
| A field in your request has no analogue in the endpoint's dialect | 501 translation_unavailable |
| This build cannot rewrite this dialect pair's responses | 501 translation_unavailable — refused before the upstream is dialled, while a refusal is still free |
| Translation was granted and requested but your dialect was undetermined | 501 translation_unavailable |
| Your body does not parse as a request of your own dialect | 400 invalid_request |
The last row is the one to internalise. 501 and 400 are split deliberately: "fix your
body" and "wait for a deploy" are different remedies, and reporting the first as the
second sends you to ask for a release you do not need.
All four translation_unavailable arms are a governance allow. Every gate said yes and
the facade could not serve the request, so the ledger records an allow and no audit hop
fires — a hash-chained record of a governance denial that never happened is worse than no
record. 501 also carries no class, and that is a ruling rather than an omission:
unavailable would promise a retry may succeed, which is false until a different binary
ships.
501is not only translation, since 2026-09-08.unsupported_dialectis served there too — see §9. Both are "this build cannot serve it", which is what the status and the absentclasssay. They differ in the ledger: the four arms above are always an allow;unsupported_dialectis adenywhen the resolver's credential gate refuses it and anallowwhen the handler does. So do not read501as "never in the audit chain" — readerror.code.
Mid-stream is different, and you should know before you see it. Once the first frame is out the status line is committed, so a response frame with no analogue cannot be refused. It is reported in band — an error event in your own dialect — and the rewriter stops emitting. Dropping it or relaying it untranslated are both worse.
And know what you will not see: the ledger cannot tell you it happened. A mid-stream
translation_unavailable writes a usage row indistinguishable from a complete call on
every ledger-facing field — status_code, termination_reason, client_gone and
inspected are all identical to the clean stream (#3134, measured; open,
prod-enablement). You are also metered for the upstream's full token count, correctly,
because the observer sees upstream frames by design. So the in-band error event is the
only record that half your answer was rewritten and half was not — if you are
reconciling spend against completions, do not expect the ledger to flag these.
Translation is not a metering variable. The usage observer is fed the upstream frame while you are fed the rewritten one, from the same line, so a translated call meters byte-identically to the same upstream stream on the pass-through path.
5. Why you were refused — the vocabulary
Every refusal carries a stable machine code. Key on error.code; never on the message.
{"error":{"code":"budget_exceeded_team","message":"…","gate":"7d/budget","class":"quota"}}
code— the stable identity. It is also what lands inllm_usage_events.deciding_gate, so it is the string you group by in a dashboard.gate— the pipeline step that refused, e.g.7d/budget. Omitted when no pipeline step refused; absence is meaningful.class— what the status means for your retry logic. Omitted at a status this build has ruled unclassifiable (500,501); absence there is a declared ruling, not a gap. Both mean the same thing — this build cannot serve it, and no retry succeeds until a different binary is deployed — and there is deliberately no class that says so (#3245 asks whether there should be).
Four classes, and the difference is load-bearing:
| Class | Statuses | What it tells your client |
|---|---|---|
quota | 402, 429 | You exhausted a cap. Retrying now cannot succeed; retrying after the window rolls, or after the cap is raised, can |
policy | 401, 403 | A governance decision refused this call. Retrying cannot succeed at any time — someone has to change a configuration |
request | 400, 404, 413 | The request is the problem |
unavailable | 502, 503 | The control could not be evaluated. You did nothing wrong and a retry may well succeed |
403 and 503 must never be collapsed. A 403 sends an operator to the policy; a 503 sends
them to the dependency. Every gate's fail-closed arm therefore has its own code
(clearance_unavailable, not insufficient_clearance) — an operator grouping by the wrong
one looks at an agent's clearance, finds it correct, and has nowhere to go.
5.1 Authentication failures are a different envelope — and this is a real wart
A refusal from the auth middleware, i.e. before the resolver runs, does not carry
code, gate or class. It carries type:
$ curl -si -XPOST http://127.0.0.1:8091/v1/chat/completions \
-H 'Authorization: Bearer notatoken' -d '{}' # verified 2026-09-07
HTTP/1.1 401 Unauthorized
{"error":{"message":"this bearer token could not be validated","type":"invalid_token"}}
$ … -H 'Authorization: Bearer uq_key_nope' # verified 2026-09-07
{"error":{"message":"invalid or expired API key","type":"invalid_api_key"}}
$ … with no Authorization header at all # verified 2026-09-07
{"error":{"message":"missing or malformed Authorization header","type":"invalid_api_key"}}
The two 401 bodies share a shape deliberately — a client that must parse two envelopes will
parse one of them wrong — and type is the field you switch on. The contract's
unauthenticated code (401) is the resolver's arm: it is reached only when the
middleware admitted a request and no caller resolved, which is a wiring fault.
Note the dispatch rule: a bearer value not beginning uq_key_ goes to the OAuth
path; anything else — including no header at all — goes to the API-key path. That is why a
missing header reports invalid_api_key.
MG-3.6 asks that authentication failure be machine-distinguishable from the other families,
and it is (type vs code, and the two are never both present). But it is
distinguishable by a different field, which is a rougher edge than the rest of the
contract and is flagged rather than papered over.
5.2 The full refusal contract
One declaration (internal/modelgateway/errorcontract.go), one status per identity, keyed
by reason — denial and refuse do not take a status at all, they look it up. The table
below is generated from that declaration and pinned to it by
test/lint/model_gateway_error_contract_doc_test.go: adding a refusal without adding a row
here fails the build, and a row here naming a code the gateway cannot emit fails it too.
ledger is what llm_usage_events.governance_outcome records and, inseparably, whether
the hash-chained agent_audit_log hop fires — a deny row and an audit row are written
together or not at all. allow means governance permitted the call and something after
it failed; recording those as denials would put a decision that was never taken into a
tamper-evident chain.
| Code | Status | Class | Ledger | Retry-After |
|---|---|---|---|---|
agent_ambiguous | 403 | policy | deny | — |
agent_credential_required | 403 | policy | deny | — |
agent_enablement_malformed | 403 | policy | deny | — |
agent_enablement_unavailable | 503 | unavailable | deny | — |
agent_not_found | 403 | policy | deny | — |
agent_scoped_key_demoted | 403 | policy | deny | — |
agent_unresolved | 403 | policy | deny | — |
auth_mode_unknown | 403 | policy | deny | — |
binding_kind_unknown | 403 | policy | deny | — |
binding_unit_corrupt | 403 | policy | deny | — |
binding_unit_unresolved | 403 | policy | deny | — |
binding_withheld | 403 | policy | deny | — |
budget_exceeded_agent | 402 | quota | deny | — |
budget_exceeded_org | 402 | quota | deny | — |
budget_exceeded_team | 402 | quota | deny | — |
budget_unavailable | 503 | unavailable | deny | — |
capability_unconfigured | 403 | policy | deny | — |
clearance_unavailable | 503 | unavailable | deny | — |
content_filter | 400 | request | deny | — |
credential_missing | 403 | policy | deny | — |
credential_store_unbound | 503 | unavailable | deny | — |
dialect_ambiguous | 400 | request | deny | — |
dialect_mismatch | 400 | request | deny | — |
egress_destination_refused | 403 | policy | allow | — |
endpoint_disabled | 403 | policy | deny | — |
endpoint_not_enabled_for_agent | 403 | policy | deny | — |
endpoint_pending_approval | 403 | policy | deny | — |
endpoint_unmetered | 403 | policy | deny | — |
governance_backend_unavailable | 503 | unavailable | deny | — |
guardrail_input_unreadable | 400 | request | deny | — |
insufficient_clearance | 403 | policy | deny | — |
invalid_request | 400 | request | deny or allow | — |
model_allowlist_unavailable | 503 | unavailable | deny | — |
model_not_enabled_for_agent | 403 | policy | deny | — |
model_not_in_team_allowlist | 403 | policy | deny | — |
model_not_served_by_endpoint | 403 | policy | deny | — |
model_required | 400 | request | deny | — |
no_approved_binding | 403 | policy | deny | — |
no_matching_endpoint | 404 | request | deny | — |
rate_limit_unavailable | 503 | unavailable | deny | — |
rate_limited_agent | 429 | quota | deny | required |
rate_limited_org | 429 | quota | deny | required |
rate_limited_team | 429 | quota | deny | required |
refusal_uncontracted | 503 | unavailable | deny | — |
request_body_too_large | 413 | request | deny or allow | — |
route_unresolved | 502 | unavailable | unassessed | — |
service_agent_budget_required | 403 | policy | deny | — |
service_governance_unavailable | 503 | unavailable | deny | — |
service_policy_required | 403 | policy | deny | — |
streaming_unsupported | 500 | (no class — declared omission) | allow | — |
translate_opt_in_malformed | 400 | request | deny | — |
translation_not_permitted | 403 | policy | deny | — |
translation_unavailable | 501 | (no class — declared omission) | allow | — |
unauthenticated | 401 | policy | deny | — |
unsupported_dialect | 501 | (no class — declared omission) | deny or allow | — |
upstream_unavailable | 502 | unavailable | allow | — |
5.3 The gate chain, in operator language
The order below is enforcement order, read back from the declaration at startup
(gates= on the wired line). Earlier gates win, so the code you get names the first
thing that was wrong, not the only one.
| Step | Refused because | The codes you will see | Where you fix it |
|---|---|---|---|
| 3 identity | No caller resolved, or an agent-scoped uq_key_* was presented after the window closed | unauthenticated, agent_scoped_key_demoted | Re-grant an OAuth token (§3) |
| 4 rate limit | Per-identity token bucket, before the body is read | rate_limited_{org,team,agent}, rate_limit_unavailable | §6.2 |
| 5 request / routing / binding | No model in the body; no approved endpoint serves it; the endpoint is not bound to your team, is pending, or is withheld | model_required, invalid_request, request_body_too_large, no_matching_endpoint, endpoint_disabled, endpoint_pending_approval, no_approved_binding, binding_withheld, binding_unit_unresolved, binding_kind_unknown | Steps 1–4 of §1 |
| 6 mode | The request names two dialects, or names one the endpoint does not speak, or opted into translation the binding does not grant | dialect_ambiguous, dialect_mismatch, translation_not_permitted, translate_opt_in_malformed | The request, or the binding's allow_translation |
| 6 metering | The endpoint has produced results the gateway could not meter, and a cap is in force | endpoint_unmetered | Fix metering, or take the L5 allow_unmetered acknowledgement — which turns off a spend control |
| 7a clearance floor | The agent's clearance is below max(endpoint, binding).min_clearance. Resolved server-side per call, never from a token claim, so an L5 demotion takes effect on the next call | insufficient_clearance, agent_not_found, clearance_unavailable | The agent's clearance, or the floor |
| 7b agent enablement | This agent's narrow-only selection does not include this endpoint or model | endpoint_not_enabled_for_agent, model_not_enabled_for_agent, agent_unresolved, agent_ambiguous, agent_enablement_malformed, agent_enablement_unavailable | The agent's Models selection (step 5) |
| 7c model allowlist | The model is not in the team's allowlist, or the endpoint does not serve it | model_not_in_team_allowlist, model_not_served_by_endpoint, binding_unit_corrupt, model_allowlist_unavailable | The binding's model list |
| 7d budget | A cap is exhausted at org, team or agent | budget_exceeded_{org,team,agent}, budget_unavailable | §6.1 |
| 7e guardrail | Content refused, or the body could not be read as a request | content_filter, guardrail_input_unreadable | The prompt |
| 8 credential | No credential at any grain this binding permits; an empty one; a store that could not be read tenant-bound | credential_missing, agent_credential_required, credential_store_unbound, unsupported_dialect | Step 3 of §1 |
agent_unresolved vs agent_ambiguous is worth a sentence, because it looks like an
internal detail and is not. member_api_keys.scoped_agent_id is TEXT with no foreign key
and its referent is genuinely ambiguous between agents(id) and agent_configurations(id).
The reader matches both and reports a value resolving two different rows as its own
refusal rather than picking one — because on a narrowing control, picking wrong misses in
the fail-open direction.
6. Quotas — what is actually enforced
6.1 Budget — 402
Three grains — org, team, agent — are all loaded and composed min-of: the
most restrictive wins, and the refusal names the deciding grain in its code
(budget_exceeded_team). An agent with a low cap cannot borrow team headroom; that is
the correct default, since a runaway agent is the failure mode caps exist to bound.
The check is pre-call. The call in flight when the cap was reached completes; the next one is refused. There is no mid-stream kill — that would bill you for a truncated answer to a question you did ask.
State the cap precisely. As of T14 (#2916, migration 223, closing #2861) and
the window unification that closed #2928 (#3126 merged 2026-09-07T10:24:59Z,
#3119 merged 13:55:22Z, #2928 closed 13:55:24Z, migration 235):
| Claim | True? |
|---|---|
| A monthly, per-period token cap whose window advances without human action, at all three grains | Yes. The window is derived — billing_period_bounds(billing_anchor_at, now()) inside the budget statement — so period_start <= now() < period_end holds by construction. Nothing has to advance it, and no background job's outage can disable it. Since #2928 there is one WITH period AS (…) CTE and org_period_start / team_period_start / agent_period_start all read it, so a min() across grains compares intervals that actually coincide |
| A money cap | No. The unit is tokens off llm_usage_events. cost_usd is deliberately NULLable (unknown is not $0), so a dollar cap is not computable from this ledger. The constant is named OrgTokenCapEnforced for exactly this reason (#3062 ruling (c)) |
| The same guarantee at team and agent grain | Yes, since 2026-09-07 — this row said "No" until that afternoon. Product ruled RECUR on 2026-09-03 and internal/modelgateway/budgetsource.go now opens the section with "THE RESIDUAL IS CLOSED: ALL THREE GRAINS NOW SHARE ONE DERIVED WINDOW (#2928)". The team and agent window predicates were deleted because a total derivation made them inert, and migration 235 marks the stored period_start / period_end columns vestigial and NULLABLE — so a surviving reader of the stored window now drops every cap, which is why they went rather than being left as a live-looking check |
The hazard this table used to name, now closed. It was worse than "no cap": under min-of composition a lapsed team row did not withdraw an opinion, it promoted the team to the org cap — the control widened, in the fail-open direction, on a timer, with nothing denying, so the refusal that names the deciding grain never fired precisely when a cap had stopped existing, and alert_pct / throttle_pct hung off the same row so the warning lapsed with the cap it was meant to warn about. Recorded rather than deleted: if you are reading a ledger from before 2026-09-07, this is what you are looking at | |
| "This tenant's spend is capped", as a composition | Still not writable — but settled, not open. #3062 closed 2026-09-07 under ruling (c): the constant was renamed to stop over-claiming, rather than the composition being completed. So the composed sentence is as false as it was; what changed is that it is now a recorded naming decision instead of an unresolved concern. Do not write it |
The AT LEAST wording on a 402 is not hedging. A window containing calls whose
metering_confidence is not complete has a used that is a floor, not a total, so
the refusal says so and counts the blind calls:
AT LEAST 812340 of 1000000 tokens used — 3 call(s) in this window carry no complete metering, so the true figure is higher
A precise number would send you looking for a billing bug. The same floor applies in the admitting direction, where it costs money rather than denying — that branch logs the incompleteness rather than staying silent (#3059).
Also enforced since #3118: all four token classes are summed (input, output,
cache_creation, cache_read). Two of them were invisible to the cap before.
6.2 Rate limits — 429
Per-identity token buckets in Redis, evaluated at step 4, before the body is read — a limiter running after a 32 MiB read would already have paid the cost it exists to bound.
| Class | Sustained | Burst | Applies to |
|---|---|---|---|
| org | 200/s | 400 | every caller — this class is total |
| team | 100/s | 200 | unit-scoped credentials |
| agent | 20/s | 40 | agent-scoped credentials |
Verified off the running process (rate_limits= on the wired line). There is no
environment variable that changes these — deliberately, so the running value and the
stated reason cannot drift. They are org × 1, × ½, × ⅒, with burst × 2, and the
ratios are containment statements: a team may not out-consume its org, an agent may not
out-consume its team.
The team class is deliberately oversubscribed — three teams at org/2 sum to 150% of the org rate. The per-team bucket is a blast-radius bound on one team, not a share of a fixed pie; the org bucket is what bounds the total. Partitioning instead would strand capacity whenever a team is idle and would silently change every team's limit the day a new team is created.
A caller the org class cannot bucket — no org_id — is refused, not admitted. A
non-empty bucket set is not the same as a bounded caller.
Two 429 properties your client can rely on:
Retry-Afteris always present on a 429 and never on any other refusal. That is the contract, not a convention:RetryAfterRequiredis declared per identity, and a header omitted at zero seconds would otherwise be indistinguishable from one never set.rate_limit_unavailable(503) is notrate_limited_*(429). A limit that could not be read is not a limit that was exceeded. There is a grace window (MODEL_GATEWAY_RATE_LIMIT_GRACE_SECONDS, default 5m) during which a per-replica fallback stands in for the shared bucket; after it, step 4 fails closed.
7. Bedrock — the interim workaround
Native AWS Bedrock is not supported and is not coming in v1 (PRD #2644 D3). This
is not a gap in the routing layer; SigV4 request signing is a different auth axis and does
not fit the registry's url + type + key model. Confirmed in the tree: git grep -i 'bedrock\|sigv4' returns no implementation — only ADR prose and one test fixture whose
model id happens to contain the word.
The gateway speaks exactly two dialects, openai_compatible and anthropic
(dialects_relayable on the wired line). So the interim path is to put something in front
of Bedrock that speaks one of them.
Two options. Both are the tenant's infrastructure, not UpsQuad's.
Option A — AWS's own OpenAI-compatible endpoint
If your account has an OpenAI-compatible chat-completions endpoint fronting Bedrock, register it directly:
dialect:openai_compatiblebase_url: the endpoint's publichttpsorigin- credential: the endpoint's bearer key, written through
SetLLMEndpointCredential
The gateway sends Authorization: Bearer <secret> and joins your request path onto
base_url, so a deployment prefix in the URL works.
Option B — a LiteLLM shim you host
Run LiteLLM configured with a Bedrock model, exposing /v1/chat/completions (or
/v1/messages for the anthropic dialect), and register the shim's URL as the
endpoint. LiteLLM holds the AWS credentials; UpsQuad holds a static bearer for the shim.
The constraint that catches everyone: the shim must be publicly resolvable over
https. SSRF fence rule 4 refuses any host resolving into the private/loopback/ link-local/CGNAT/ULA set unconditionally, andMODEL_GATEWAY_DEV_HTTP_HOSTS— the allowlist that would exempt one — relaxes only the scheme check and is currently inert. A shim on10.0.0.0/8, on a Kubernetes service name, or onlocalhostcannot be registered, and this is exactly why the devbox stack cannot relay to its ownupsquad-litellmcontainer.
infra/litellm/config.yaml is a working LiteLLM proxy config in this repo, but it is a
dev-only Anthropic facade over local Ollama for the Claude-SDK worker — it is not a
Bedrock configuration and it is not reachable by the Model Gateway for the reason above.
Read it as a syntax reference only.
What you give up either way, and it should be said before someone finds out from a
bill: metering fidelity is whatever the shim reports. If it does not return a usage block
the extractor recognises, calls land as metering_confidence not-complete, and every
budget number over that window becomes an AT LEAST floor (§6.1) —
including the one that decides whether to admit the next call.
8. Where the spend shows up
Tenant-facing, per org/team/agent/model — AIGatewayUsageService, read-only by
construction (T16 #2918):
| RPC | Returns |
|---|---|
QueryUsage | Usage at org / team / agent / model / endpoint / day / session grain, clearance-gated and paginated |
ExportUsage | The same as CSV |
QueryLatency | Percentiles that state the population they were computed over (p99_of_measured_ms) |
GetBudgetBurndown | The cap in force at one grain, and a projection that refuses or ranges whenever a cost is unknown |
unknown cost is a separate bucket, never summed into the total and never dropped
(OQ-3). A total that silently absorbed unpriced calls would be confidently wrong in the
flattering direction on exactly the largest ones.
The five dashboards (MG-3.7) are upsquad-client work and are all delivered as of
2026-09-07: Spend + CSV/API export (client#840, merged 2026-09-05T16:24:21Z); Agent
Economics, Reliability, Governance and Budget burndown (client#842, merged
2026-09-07T13:39:49Z). Both frontends carry every surface (two-frontend parity). This
paragraph said the second four were "open on client#842" until that PR merged — client#842
is a merged PR, not an open issue, and MG-3.7 is complete.
Platform-operator view — Grafana, Model Gateway — Overview
(deployments/observability/grafana/model-gateway-overview.json). Panels as of #3011/#3008:
concurrency and in-flight upstream requests (the HPA's own input), scrape health, peak
concurrency per replica, rate-limit refusals by identity class, rate-limit decisions by
arm, step-4 degradation, budget refusals by deciding grain, budget decisions by arm,
gate 7d failing closed, and unclassified gate-7d decisions.
Two panels exist because absence had to become visible: "gate 7d failing closed" separates a control that denied from one that could not be evaluated, and "unclassified" catches a future branch that classifies nothing rather than letting it go silent.
Denials vs allows are recorded differently, on purpose. A denial writes a
llm_usage_events deny row and a hash-chained agent_audit_log hop. An allowed call
writes a ledger row and no audit hop — chaining every served call would put an
advisory-lock round trip on the hot path (PRD v1.2 amendment to MG-2.11).
9. What this page does not claim
-
The hostname being live is not production traffic. The flip released E2E testing traffic under the scoped waiver on #2912; production tenant LLM traffic is still gated. #2912 itself stays open — its closure is the founder's call given the waiver scope, not a consequence of the flip.
Do not read a gate list off this page. The register moves faster than a runbook can. On the afternoon of 2026-09-07 three things this page asserted as open closed inside eighteen minutes —
client#842merged13:39:49Z, #2928 closed13:55:24Z, #3062 closed13:57:24Z. The pre-promotion checklist is a query:gh api "search/issues?q=repo:upsquad-ai/upsquad-core+label:prod-enablement+is:open&per_page=100" \--jq '.items[] | "#\(.number) \(.title)"'Filter it to the offering being shipped, per the Approval Protocol in
CLAUDE.md. At the time of writing, #3120 is the gate whose title names this switch directly.And re-sweep the whole page, not the row you came for. Every state claim here carries a date for exactly this reason; the ones that rot are never the ones you are looking at:
grep -oE '#[0-9]{3,5}' docs/runbooks/model-gateway-external-agent-onboarding.md | tr -d '#' | sort -un |while read -r n; dogh api repos/upsquad-ai/upsquad-core/issues/"$n" \--jq '"#\(.number) [\(.state)] closed=\(.closed_at // "-") \(.title[0:60])"'doneCross-repo references (
client#…) need the same check against their repo — and check whether the number is an issue or a merged PR, because "open on client#842" was wrong in both halves. -
The
:8091HA posture is not established. #2646 remains an openprod-enablementgate that slice 3 does not close. One devbox host edge, one process. -
GKE ingress for the Model Gateway is undesigned.
infra/edge/envoy.yamlgoverns the devbox host edge only; the cluster Service isClusterIPwith no Ingress, no HTTPRoute, no LoadBalancer and no ArgoCD Application, and an exposure check asserts that stays true. #2987 — deliberately notprod-enablement, because it gates no switch that exists.One thing that check no longer catches, recorded once so nobody rediscovers it as a surprise.
scripts/check-model-gateway-exposure.pyfinds edge references through threeEDGE_MATCHERSarms: the tokenmodel-gateway, the tokenmodel_gateway, and a digit-bounded8091. The public hostname used to satisfy the first arm by beingmodel-gateway.app.upsquad.ai; after the #3196 / #3198 renames it ismodel-gw-beta.upsquad.ai, which contains nomodel-gatewaytoken. Coverage is unchanged in practice — the vhost'sname:field is stillmodel-gateway, and anything that actually reaches the gateway must carrymodel_gateway(the cluster) or8091(the port), so both other arms still fire and the census is unchanged at 4 objects. But the hostname string is no longer load-bearing for that arm. Do not prune an arm on the reasoning that "the hostname covers it". -
The edge authn/authz/rate-limit baseline is #2985, ruled and re-scoped. Ruled option (b), gateway-enforced (founder, 2026-09-07) — that released the flip, and #2985 no longer gates it. It stays open scoped to first external-tenant onboarding. The consequence to carry: under (b) there is no Access application in front of the hostname, so the
direct_responseseal in §2 is the sole control on the unauthenticated ops surface. Any future edit to that vhost is a security change, not a routing change. -
unsupported_dialectmoved 502 → 501, and if you wrote a client before 2026-09-08 this is a breaking change to one branch. Founder ruling on the #2981 rider, 2026-09-08. It was served at 502, which the taxonomy maps tounavailable— "the control could not be evaluated; a retry may well succeed" — and for this refusal every control that ran said yes and no retry can succeed until a binary with a transport for the dialect ships. That is verbatim the sentence the 501 ruling declares false for the identical fact, so the code was telling clients to back off and retry a permanent condition.What changes on the wire: the status, and the
classfield, which is now absent — 501 is a declared-unclassified status, so the refusal carries no class rather than a wrong one. Nothing else. Thecodeis unchanged, so a client keying onerror.code(which is what §5 asks you to do) needs no change at all; a client branching on the status or onclass == "unavailable"does.Why now rather than never, or later. The only public-hostname traffic today is E2E waiver traffic, so this costs its minimum; after first external-tenant onboarding it would have needed deprecation ceremony for a code nobody could have been relying on correctly.
Its ledger outcome did not move and is still per-producer: refused pre-admission at gate
8/credentialit is adenyand is in the audit chain; refused post-admission by the handler it is anallowwith no hop. The class is the same at both and the disposition is not, which is not an inconsistency — the class answers "could a retry ever succeed" (no, either way) and the disposition answers "did a gate decide". Two questions.With that ruled, no contract disposition is contested any longer; the register that tracked them is deleted rather than left empty. One question it raises is open and tracked at #3245: the 501 ruling declines to invent a fifth
classand stated it would be revisited if a second refusal in this shape appeared. Three now have. That is a vocabulary question, not a defect — the current state is a declared omission and is honest on the wire — but it has a closing window for the same reason this change had one. -
The ledger half of #2981 is fixed, and it changes what you will find in the chain. Option B of the #2981 fork (founder ruling, 2026-09-07 — not to be confused with #2985's option (b) above, ruled the same day) moved the disposition from the refusal reason onto the producer. If you audited this gateway before that landed, two things are now different and both are corrections:
route_unresolvedused to writegovernance_outcome='deny', adeciding_gateand a hash-chained audit hop for a request where no gate ever decided. It now writes a row withgovernance_outcomeNULL — the ledger's spelling for "no gate chain evaluated this call" — and no audit hop.invalid_request,request_body_too_largeandunsupported_dialectare each emitted by two producers. Refused pre-admission by the pipeline they are denials and are in the chain, as before. Refused after every gate said yes they now record an allow and produce no hop, because governance permitted those calls. Thedeny or allowin the table above is that fact, per request.
What this costs you as an auditor, stated plainly: the refusal reason for a non-deny row is not in
deciding_gate— migration 207 permits that column only beside a deny — so it is counted onmodel_gateway_refusal_ledger_disposition_total{reason,disposition}and nowhere durable. That metric has a retention horizon and is not tamper-evident. The trade was made deliberately: what the chain now contains is only decisions that happened. Expect a one-off step down ingovernance_assessed_deny_eventsat the deploy and a matching step up ingovernance_unassessed_events; alert rules keyed on either need the note. That step is now observable rather than hypothetical: the flip released E2E testing traffic (first bullet), so these counters have a live population to move. -
Nothing on this page was executed end-to-end through a successful governed call, for the reason in §3: the devbox has no registered external client and no registered endpoint. The flip did not change this — a live hostname that answers
401is still a refusal path, and a 401 from the internet is not evidence of a completed governed call. Refusal shapes, the token endpoint's resource contract, the wired-controls line, the seal battery and the DNS/TLS state are all live measurements taken 2026-09-07; success shapes are read from source.
Related
- Platform Token Issuer (MCP trust root) — the
/oauth/tokensurface, both client-authentication methods, the enable gate, the hardening posture - Completion-path provider posture, per deployment shape — a different seam (the platform's own completion path) with a different trust class; its "an RFC1918 address is fine" rule does not apply to the Model Gateway registry
- Landing a PR
- PRD #2644 · HLD #2649 · ADR-0034 #2643 · LLD #2901 · tracker #2648