Skip to main content

Connecting an external agent to the Model Gateway

Operator onboarding — MG-3.9. How an agent that runs outside UpsQuad obtains a platform token and completes a governed LLM call through the Model Gateway, what it will be refused for, and where the resulting spend shows up.

Read this before you read anything else​

The hostname is LIVE. Production tenant LLM traffic is still gated. Those are two different statements and the difference is the whole of this box.

https://model-gw-beta.upsquad.ai resolves, serves and is sealed, verified from the public internet on 2026-09-07 — the battery is in §2 and on the flip comment. What it is released for is E2E testing traffic under the scoped waiver recorded on #2912; production tenant LLM traffic remains gated — #3120 is the named open gate, and the standing list is the query in §9, not a copy on this page. Do not read "the endpoint answers" as "the platform is carrying tenant load" (LLD #2901 §0.1).

The hostname took two renames to get here and the reason is worth one line, because it constrains the next one: model-gateway.app.upsquad.ai → model-gw-beta.app.upsquad.ai (PR #3196) → model-gw-beta.upsquad.ai (PR #3198). A two-label name needs a per-hostname certificate — Cloudflare's universal cert is *.upsquad.ai + upsquad.ai, and a TLS wildcard matches exactly one label — so with Total TLS off the cert never issued. The one-label name was covered instantly. Any future MG hostname must stay one-label under upsquad.ai unless Total TLS is turned on first.

So: §2 has a devbox / loopback half and a public hostname half, and both are now live measurements rather than one measurement and one plan. Every control below names the task that delivered it, so you can re-verify it rather than re-trust this page.


1. What you are connecting, and who does what​

An external agent is a first-class agents.id — team, clearance, per-agent model enablement, budget and audit all apply to it identically to a platform agent. The only difference is the runtime = "external" provenance stamped on its token and carried onto every ledger and audit row (internal/auth/oauth/token_endpoint.go, MG-3.3 / T5 #2907).

There is no second registration concept for LLM access. The same external_agent_clients row (migration 155, #1956) that authenticates a BYO agent to the governed MCP gateway authenticates it to the Model Gateway; only the requested audience differs (MG-3.3).

Onboarding is five writes, and they are deliberately split across three authorities:

#StepSurfaceWhoDelivered by
1Register the upstream endpointAIGatewayEndpointRegistryService.RegisterLLMEndpointL4slice 1 T9 #2685
2Activate itSetLLMEndpointStatus → approvedL5slice 1 T9 #2685
3Write the provider credentialAIGatewayEndpointCredentialService.SetLLMEndpointCredentialL4/L5 per grainslice 1 T10 #2686
4Bind the endpoint to a team, then approve the bindingCreateLLMBinding → ApproveLLMBindingL4 + the unit's Managerslice 2 T2 #2807
5Register the external agent + narrow its model accessAgentService.RegisterExternalAgentL4#1956 + slice 2 T4 #2810

Why registration and activation are different clearances. base_url is an egress destination. Authoring one is L4; turning it into a destination the platform will dial on your behalf is an L5 decision. This split is normative and is not relaxable to auto-approve (proto/upsquad/aigateway/v1/endpoint.proto).

Why binding approval is L4-and-Manager rather than L5. A binding grants an already-approved endpoint to one unit, and the authority over a unit is its Manager (org_units.manager_member_id). A unit with no Manager fails closed — there is no authority to approve, so nobody may.

1.1 The endpoint URL fence — read this before you pick a base_url​

base_url is tenant-supplied, so it goes through the full SSRF fence (internal/modelgateway/registry/baseurl.go, LLD #2673 §8.3). Four rules, applied in order; every refusal message begins base_url rule N: so you can tell which one rejected the value.

RuleRequirement
1Absolute http/https URL with a host
2https only. Plaintext http is permitted solely for hosts in MODEL_GATEWAY_DEV_HTTP_HOSTS, which is empty in production
3No userinfo component, no fragment
4No private, loopback, link-local, CGNAT or ULA destination — unconditionally. Every A/AAAA record is tested; one blocked record refuses the URL. A host that cannot be resolved is refused, because unproven must not read as allowed

Rule 4 has no escape hatch. MODEL_GATEWAY_DEV_HTTP_HOSTS relaxes the scheme check and nothing else; it is currently inert and setting it changes nothing (docker-compose.dev.yml, LLD §8.3 v1.8). This is the single most common surprise when standing up a shim — see §7.

Not closed in v1, and stated rather than assumed away: DNS rebinding between validation and dial. A hostname that resolves public here can resolve to loopback when the gateway actually dials it. Closing it needs the pinned-resolver work the MCP egress ADR owns (LLD #2673 §13).


2. Reaching the gateway​

Today — beta-dev / the devbox stack · verified 2026-09-07​

The gateway process runs and serves, but only on loopback. There is no edge route.

WhatWhereVerified
Model Gatewayhttp://127.0.0.1:8091 on the devbox — loopback only, no host publishyes
Token endpointhttp://127.0.0.1:8083/oauth/token (the context-engine container's public mux)yes
AS metadatahttp://127.0.0.1:8083/.well-known/oauth-authorization-serveryes
JWKShttp://127.0.0.1:8083/.well-known/jwks.jsonyes
LLM audience in forcehttp://context-engine:8080/llmyes — read off the process

Get there over SSH to the devbox (docs/runbooks/devbox-ssh-warp.md). beta-dev.app.upsquad.ai is a different vhost — the portal — and it sits behind Cloudflare Access; an unauthenticated request to it returns 302, not the gateway.

The process states its whole posture in one startup line. This is the artefact to read before debugging anything, and it is the source of the numbers on this page:

$ docker logs upsquad-model-gateway | grep -m1 'controls wired'
… gates="7a/clearance_floor -> 7b/agent_enablement -> 7c/model_allowlist -> 7d/budget -> 7e/guardrail"
rate_limits="org 200/s burst 400 -> team 100/s burst 200 -> agent 20/s burst 40"
dialects_relayable="[anthropic openai_compatible]" credentials_enabled=false
org_token_cap_enforced=true production_routing_released=true
agent_token_auth=true agent_token_audience=http://context-engine:8080/llm
agent_scoped_keys=refused environment=development

Read it off the process you are talking to, not off this page. The excerpt above was taken from the devbox container on 2026-09-07 and the fields quoted are stable across builds, but the controls= and seam_types= lists grow with each task and the devbox image lags main. A build's own line is authoritative for that build; this one is an example of where to look.

Two lines of that are load-bearing for anyone testing here:

  • credentials_enabled=false — the devbox has no credential store bound, so a call that passes every gate is refused at step 8 with credential_store_unbound (503). That is the stack, not your configuration.
  • agent_scoped_keys=refused — the uq_key_* deprecation window is closed on this stack (T17 #2919, merged #3068). An agent-scoped API key is refused with agent_scoped_key_demoted (403). OAuth is the only agent path here.

The public hostname — LIVE, verified 2026-09-07​

WhatWhereState
Model Gatewayhttps://model-gw-beta.upsquad.ai/v1/…live, edge installed at 2ec6629e3 (#2912)
Sealed bycatch-all direct_response: 404 — only /v1/ is routedlive, probed
Released forE2E testing traffic, scoped waiver on #2912production tenant traffic still gated (#3120)

The vhost routes only /v1/ and seals everything else. That seal is the point: cmd/model-gateway's mux serves GET /metrics, /readyz and /healthz with no authentication on the same listener, so a prefix: "/" route would have published all three. Under #2985's ruling — option (b), gateway-enforced, so there is no Cloudflare Access application in front — that seal is not defence in depth. It is the only thing between those three handlers and the internet.

So it was probed rather than asserted. From the public internet, 2026-09-07:

POST /v1/chat/completions (no auth) -> 401 {"error":{"message":"missing or malformed
POST /v1/chat/completions (Bearer nope) -> 401 Authorization header","type":"invalid_api_key"}}
GET /metrics -> 404 ┐
GET /readyz -> 404 │ the direct_response seal
GET /healthz -> 404 │
GET / -> 404 ┘
beta-dev.app.upsquad.ai -> 302 · gw-dev.upsquad.ai -> 302 neighbours unaffected

Two details make that battery mean something, and a bare status code would not.

  1. The 404s carry no x-envoy-upstream-service-time header; the 401 does. So the 401 is the gateway refusing through Envoy, and the four 404s never reached an upstream at all — they are Envoy's direct_response. A status code alone cannot tell "the seal fired" from "the gateway happened to 404".
  2. The same /metrics is 200 on loopback (curl 127.0.0.1:8091/metrics). The seal is measurably converting a live 200 into a 404, which is the claim; four 404s on their own are equally consistent with a broken config.

The 401 body is the gateway's own JSON and there is no Access interstitial, no redirect — the identity decision is being taken by T4 #2906's middleware (cmd/model-gateway/main.go), which is the mechanism ruling (b) selected, observed working from outside.

Rollback is three independent one-minute levers (any one of them takes the hostname down): delete the cloudflared ingress block and restart cloudflared-upsquad; delete the CNAME; restore /etc/envoy/backups/envoy.yaml.<TS>.bak and restart envoy. The operational checklist lives on #2912 and is not restated here — a restatement is a copy, and the copy is the half that goes stale.

The trap: the audience is not derived from the gateway's hostname​

LLMAudienceURI(publicBaseURL) = publicBaseURL + "/llm", where publicBaseURL is the platform base URL (MCP_GATEWAY_PUBLIC_BASE_URL) — the same source TeamAudienceURI and PlatformAudienceURI use, and distinct from both so an MCP token can never satisfy the Model Gateway (LLD §3.5, T4 #2906).

Requesting resource=https://model-gw-beta.upsquad.ai/llm does not work — not before the flip, not now that the hostname is live, and the two renames it took to get to this name are the argument. Deriving aud from the gateway's hostname was considered and refused: the minter would then need to know that hostname, and a hostname change would silently invalidate every outstanding token with a 401 that names nothing. That hypothetical happened twice in one day (#3196, #3198) and cost this contract nothing, which is the design working.

Derive it, never guess it: read issuer from the authorization server's metadata and append /llm. That is the same derivation the validator performs.


3. Getting a token · verified 2026-09-07​

grant_type=client_credentials only. No refresh tokens — the agent re-grants. TTL is ≈10 minutes. Two client-authentication methods, both advertised in the AS metadata: private_key_jwt (RFC 7523, preferred — you hold the private key, we store only the public JWK) and client_secret_post (fallback, bcrypt).

The mechanics of both, the hardening posture (uniform 401 invalid_client for every failure, per-IP and per-client rate limits, the external_token_issued / external_token_denied audit rows) and the MCP_GATEWAY_ENABLED / MCP_GATEWAY_TENANTS enable gate are documented once, in Platform Token Issuer. Everything there applies unchanged. What follows is the one thing that differs for the Model Gateway.

The resource parameter is the whole difference​

BASE=http://127.0.0.1:8083 # the PLATFORM origin, not the gateway's

# Derive the LLM audience from the issuer — never hand-write it.
LLM_AUD="$(curl -s "$BASE/.well-known/oauth-authorization-server" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["issuer"].rstrip("/")+"/llm")')"

curl -s -X POST "$BASE/oauth/token" \
-d grant_type=client_credentials \
-d "client_id=$CLIENT_ID" -d "client_secret=$CLIENT_SECRET" \
-d "resource=$LLM_AUD"
# => {"access_token":"<jwt>","token_type":"Bearer","expires_in":600}

RFC 8707 resource selects which audience is minted. It is matched exact, never by prefix — a prefix match would let …/llm-evil select the LLM audience.

resourceResultVerified
absentMCP-gateway audience (the pre-#2906 shape; every existing client is unaffected)yes — 401 invalid_client for a bogus client, i.e. the resource was accepted
the exact LLM audienceModel Gateway audienceyes — 401 invalid_client, i.e. accepted
anything else400 invalid_target — "the requested resource is not served by this authorization server"yes

An unrecognised resource is refused rather than ignored. Ignoring it would hand back an MCP-audience token to a client that asked for the LLM one, which then fails one hop away as a 401 that can only say "wrong audience".

What was NOT executed here. The devbox has zero rows in external_agent_clients and zero in llm_endpoints, so no successful mint and no successful relay were performed. Every refusal shape above is a real response from the running stack; the success shapes are read from the source and the Platform Token Issuer runbook, and are marked as such.


4. Making the call​

# PUBLIC (live, verified 2026-09-07): MG=https://model-gw-beta.upsquad.ai
# DEVBOX LOOPBACK (unsealed — serves /metrics, /readyz, /healthz):
# MG=http://127.0.0.1:8091

curl -s -X POST "$MG/v1/chat/completions" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"model":"<a model id on an endpoint your team is bound to>",
"messages":[{"role":"user","content":"hello"}]}'

Things worth knowing before the first call:

  • Everything under /v1/ is relayed. The gateway does not enumerate upstream routes — a registered base_url may carry an Azure-style deployment prefix and your path is joined onto it. So /v1/embeddings, /v1/rerank and provider-native paths reach the upstream; a fixed route table would break the endpoints the registry exists to support.
  • Your credentials are stripped. Authorization, x-api-key, Api-Key, Cookie and every X-Upsquad-* header are removed before forward, and the tenant's credential is injected: Authorization: Bearer … for openai_compatible, x-api-key: … for anthropic (internal/modelgateway/transport/headers.go).
  • Dialect is inferred from the request shape, not from a route table (T6 #2908). A body with system / top_k / thinking names anthropic; one with frequency_penalty / response_format / seed names openai_compatible. A request naming two is refused dialect_ambiguous (400). A request naming none is not refused — it falls to pass-through, which is what every earlier build did with it.
  • Cross-dialect translation is opt-in twice. The team's binding must carry allow_translation, and the request must send X-Upsquad-Translate: 1. Never inferred: PRD §11 forbids silent translation and inference is exactly that silence. See §4.1 for what happens when it is on.
  • The credential grain walk is most-specific-wins with no fallback where the binding forbids one. If the binding sets credential_mode = per_agent with require_per_agent, there is deliberately no fallback to the team or org key (MG-2.2); the refusal is agent_credential_required.
  • HTTP/1.1 only. :8091 does not terminate h2c. An h2-terminating proxy in front is fine.

4.1 Translation: anthropic ⇄ openai_compatible​

Both directions, request and response, shipped by T7 #2909 (merged 2026-09-07). Two dialects only — no third, no partial-fidelity mode, and no fallback to pass-through on a translation failure: a failure is a refusal.

The rule that makes it honest: a request feature with no cross-dialect analogue is REFUSED, never dropped. A dropped tool_choice, thinking block or response_format produces a successful call with different semantics, which is the worst failure shape available because you are never told. The lossy set is derived from the dialect structs rather than hand-listed, so a field a vendor ships tomorrow refuses by default.

What happenedWire
A field in your request has no analogue in the endpoint's dialect501 translation_unavailable
This build cannot rewrite this dialect pair's responses501 translation_unavailable — refused before the upstream is dialled, while a refusal is still free
Translation was granted and requested but your dialect was undetermined501 translation_unavailable
Your body does not parse as a request of your own dialect400 invalid_request

The last row is the one to internalise. 501 and 400 are split deliberately: "fix your body" and "wait for a deploy" are different remedies, and reporting the first as the second sends you to ask for a release you do not need.

All four translation_unavailable arms are a governance allow. Every gate said yes and the facade could not serve the request, so the ledger records an allow and no audit hop fires — a hash-chained record of a governance denial that never happened is worse than no record. 501 also carries no class, and that is a ruling rather than an omission: unavailable would promise a retry may succeed, which is false until a different binary ships.

501 is not only translation, since 2026-09-08. unsupported_dialect is served there too — see §9. Both are "this build cannot serve it", which is what the status and the absent class say. They differ in the ledger: the four arms above are always an allow; unsupported_dialect is a deny when the resolver's credential gate refuses it and an allow when the handler does. So do not read 501 as "never in the audit chain" — read error.code.

Mid-stream is different, and you should know before you see it. Once the first frame is out the status line is committed, so a response frame with no analogue cannot be refused. It is reported in band — an error event in your own dialect — and the rewriter stops emitting. Dropping it or relaying it untranslated are both worse.

And know what you will not see: the ledger cannot tell you it happened. A mid-stream translation_unavailable writes a usage row indistinguishable from a complete call on every ledger-facing field — status_code, termination_reason, client_gone and inspected are all identical to the clean stream (#3134, measured; open, prod-enablement). You are also metered for the upstream's full token count, correctly, because the observer sees upstream frames by design. So the in-band error event is the only record that half your answer was rewritten and half was not — if you are reconciling spend against completions, do not expect the ledger to flag these.

Translation is not a metering variable. The usage observer is fed the upstream frame while you are fed the rewritten one, from the same line, so a translated call meters byte-identically to the same upstream stream on the pass-through path.


5. Why you were refused — the vocabulary​

Every refusal carries a stable machine code. Key on error.code; never on the message.

{"error":{"code":"budget_exceeded_team","message":"…","gate":"7d/budget","class":"quota"}}
  • code — the stable identity. It is also what lands in llm_usage_events.deciding_gate, so it is the string you group by in a dashboard.
  • gate — the pipeline step that refused, e.g. 7d/budget. Omitted when no pipeline step refused; absence is meaningful.
  • class — what the status means for your retry logic. Omitted at a status this build has ruled unclassifiable (500, 501); absence there is a declared ruling, not a gap. Both mean the same thing — this build cannot serve it, and no retry succeeds until a different binary is deployed — and there is deliberately no class that says so (#3245 asks whether there should be).

Four classes, and the difference is load-bearing:

ClassStatusesWhat it tells your client
quota402, 429You exhausted a cap. Retrying now cannot succeed; retrying after the window rolls, or after the cap is raised, can
policy401, 403A governance decision refused this call. Retrying cannot succeed at any time — someone has to change a configuration
request400, 404, 413The request is the problem
unavailable502, 503The control could not be evaluated. You did nothing wrong and a retry may well succeed

403 and 503 must never be collapsed. A 403 sends an operator to the policy; a 503 sends them to the dependency. Every gate's fail-closed arm therefore has its own code (clearance_unavailable, not insufficient_clearance) — an operator grouping by the wrong one looks at an agent's clearance, finds it correct, and has nowhere to go.

5.1 Authentication failures are a different envelope — and this is a real wart​

A refusal from the auth middleware, i.e. before the resolver runs, does not carry code, gate or class. It carries type:

$ curl -si -XPOST http://127.0.0.1:8091/v1/chat/completions \
-H 'Authorization: Bearer notatoken' -d '{}' # verified 2026-09-07
HTTP/1.1 401 Unauthorized
{"error":{"message":"this bearer token could not be validated","type":"invalid_token"}}

$ … -H 'Authorization: Bearer uq_key_nope' # verified 2026-09-07
{"error":{"message":"invalid or expired API key","type":"invalid_api_key"}}

$ … with no Authorization header at all # verified 2026-09-07
{"error":{"message":"missing or malformed Authorization header","type":"invalid_api_key"}}

The two 401 bodies share a shape deliberately — a client that must parse two envelopes will parse one of them wrong — and type is the field you switch on. The contract's unauthenticated code (401) is the resolver's arm: it is reached only when the middleware admitted a request and no caller resolved, which is a wiring fault.

Note the dispatch rule: a bearer value not beginning uq_key_ goes to the OAuth path; anything else — including no header at all — goes to the API-key path. That is why a missing header reports invalid_api_key.

MG-3.6 asks that authentication failure be machine-distinguishable from the other families, and it is (type vs code, and the two are never both present). But it is distinguishable by a different field, which is a rougher edge than the rest of the contract and is flagged rather than papered over.

5.2 The full refusal contract​

One declaration (internal/modelgateway/errorcontract.go), one status per identity, keyed by reason — denial and refuse do not take a status at all, they look it up. The table below is generated from that declaration and pinned to it by test/lint/model_gateway_error_contract_doc_test.go: adding a refusal without adding a row here fails the build, and a row here naming a code the gateway cannot emit fails it too.

ledger is what llm_usage_events.governance_outcome records and, inseparably, whether the hash-chained agent_audit_log hop fires — a deny row and an audit row are written together or not at all. allow means governance permitted the call and something after it failed; recording those as denials would put a decision that was never taken into a tamper-evident chain.

CodeStatusClassLedgerRetry-After
agent_ambiguous403policydeny—
agent_credential_required403policydeny—
agent_enablement_malformed403policydeny—
agent_enablement_unavailable503unavailabledeny—
agent_not_found403policydeny—
agent_scoped_key_demoted403policydeny—
agent_unresolved403policydeny—
auth_mode_unknown403policydeny—
binding_kind_unknown403policydeny—
binding_unit_corrupt403policydeny—
binding_unit_unresolved403policydeny—
binding_withheld403policydeny—
budget_exceeded_agent402quotadeny—
budget_exceeded_org402quotadeny—
budget_exceeded_team402quotadeny—
budget_unavailable503unavailabledeny—
capability_unconfigured403policydeny—
clearance_unavailable503unavailabledeny—
content_filter400requestdeny—
credential_missing403policydeny—
credential_store_unbound503unavailabledeny—
dialect_ambiguous400requestdeny—
dialect_mismatch400requestdeny—
egress_destination_refused403policyallow—
endpoint_disabled403policydeny—
endpoint_not_enabled_for_agent403policydeny—
endpoint_pending_approval403policydeny—
endpoint_unmetered403policydeny—
governance_backend_unavailable503unavailabledeny—
guardrail_input_unreadable400requestdeny—
insufficient_clearance403policydeny—
invalid_request400requestdeny or allow—
model_allowlist_unavailable503unavailabledeny—
model_not_enabled_for_agent403policydeny—
model_not_in_team_allowlist403policydeny—
model_not_served_by_endpoint403policydeny—
model_required400requestdeny—
no_approved_binding403policydeny—
no_matching_endpoint404requestdeny—
rate_limit_unavailable503unavailabledeny—
rate_limited_agent429quotadenyrequired
rate_limited_org429quotadenyrequired
rate_limited_team429quotadenyrequired
refusal_uncontracted503unavailabledeny—
request_body_too_large413requestdeny or allow—
route_unresolved502unavailableunassessed—
service_agent_budget_required403policydeny—
service_governance_unavailable503unavailabledeny—
service_policy_required403policydeny—
streaming_unsupported500(no class — declared omission)allow—
translate_opt_in_malformed400requestdeny—
translation_not_permitted403policydeny—
translation_unavailable501(no class — declared omission)allow—
unauthenticated401policydeny—
unsupported_dialect501(no class — declared omission)deny or allow—
upstream_unavailable502unavailableallow—

5.3 The gate chain, in operator language​

The order below is enforcement order, read back from the declaration at startup (gates= on the wired line). Earlier gates win, so the code you get names the first thing that was wrong, not the only one.

StepRefused becauseThe codes you will seeWhere you fix it
3 identityNo caller resolved, or an agent-scoped uq_key_* was presented after the window closedunauthenticated, agent_scoped_key_demotedRe-grant an OAuth token (§3)
4 rate limitPer-identity token bucket, before the body is readrate_limited_{org,team,agent}, rate_limit_unavailable§6.2
5 request / routing / bindingNo model in the body; no approved endpoint serves it; the endpoint is not bound to your team, is pending, or is withheldmodel_required, invalid_request, request_body_too_large, no_matching_endpoint, endpoint_disabled, endpoint_pending_approval, no_approved_binding, binding_withheld, binding_unit_unresolved, binding_kind_unknownSteps 1–4 of §1
6 modeThe request names two dialects, or names one the endpoint does not speak, or opted into translation the binding does not grantdialect_ambiguous, dialect_mismatch, translation_not_permitted, translate_opt_in_malformedThe request, or the binding's allow_translation
6 meteringThe endpoint has produced results the gateway could not meter, and a cap is in forceendpoint_unmeteredFix metering, or take the L5 allow_unmetered acknowledgement — which turns off a spend control
7a clearance floorThe agent's clearance is below max(endpoint, binding).min_clearance. Resolved server-side per call, never from a token claim, so an L5 demotion takes effect on the next callinsufficient_clearance, agent_not_found, clearance_unavailableThe agent's clearance, or the floor
7b agent enablementThis agent's narrow-only selection does not include this endpoint or modelendpoint_not_enabled_for_agent, model_not_enabled_for_agent, agent_unresolved, agent_ambiguous, agent_enablement_malformed, agent_enablement_unavailableThe agent's Models selection (step 5)
7c model allowlistThe model is not in the team's allowlist, or the endpoint does not serve itmodel_not_in_team_allowlist, model_not_served_by_endpoint, binding_unit_corrupt, model_allowlist_unavailableThe binding's model list
7d budgetA cap is exhausted at org, team or agentbudget_exceeded_{org,team,agent}, budget_unavailable§6.1
7e guardrailContent refused, or the body could not be read as a requestcontent_filter, guardrail_input_unreadableThe prompt
8 credentialNo credential at any grain this binding permits; an empty one; a store that could not be read tenant-boundcredential_missing, agent_credential_required, credential_store_unbound, unsupported_dialectStep 3 of §1

agent_unresolved vs agent_ambiguous is worth a sentence, because it looks like an internal detail and is not. member_api_keys.scoped_agent_id is TEXT with no foreign key and its referent is genuinely ambiguous between agents(id) and agent_configurations(id). The reader matches both and reports a value resolving two different rows as its own refusal rather than picking one — because on a narrowing control, picking wrong misses in the fail-open direction.


6. Quotas — what is actually enforced​

6.1 Budget — 402​

Three grains — org, team, agent — are all loaded and composed min-of: the most restrictive wins, and the refusal names the deciding grain in its code (budget_exceeded_team). An agent with a low cap cannot borrow team headroom; that is the correct default, since a runaway agent is the failure mode caps exist to bound.

The check is pre-call. The call in flight when the cap was reached completes; the next one is refused. There is no mid-stream kill — that would bill you for a truncated answer to a question you did ask.

State the cap precisely. As of T14 (#2916, migration 223, closing #2861) and the window unification that closed #2928 (#3126 merged 2026-09-07T10:24:59Z, #3119 merged 13:55:22Z, #2928 closed 13:55:24Z, migration 235):

ClaimTrue?
A monthly, per-period token cap whose window advances without human action, at all three grainsYes. The window is derived — billing_period_bounds(billing_anchor_at, now()) inside the budget statement — so period_start <= now() < period_end holds by construction. Nothing has to advance it, and no background job's outage can disable it. Since #2928 there is one WITH period AS (…) CTE and org_period_start / team_period_start / agent_period_start all read it, so a min() across grains compares intervals that actually coincide
A money capNo. The unit is tokens off llm_usage_events. cost_usd is deliberately NULLable (unknown is not $0), so a dollar cap is not computable from this ledger. The constant is named OrgTokenCapEnforced for exactly this reason (#3062 ruling (c))
The same guarantee at team and agent grainYes, since 2026-09-07 — this row said "No" until that afternoon. Product ruled RECUR on 2026-09-03 and internal/modelgateway/budgetsource.go now opens the section with "THE RESIDUAL IS CLOSED: ALL THREE GRAINS NOW SHARE ONE DERIVED WINDOW (#2928)". The team and agent window predicates were deleted because a total derivation made them inert, and migration 235 marks the stored period_start / period_end columns vestigial and NULLABLE — so a surviving reader of the stored window now drops every cap, which is why they went rather than being left as a live-looking check
A lapsed team window means "no cap at that grain"The hazard this table used to name, now closed. It was worse than "no cap": under min-of composition a lapsed team row did not withdraw an opinion, it promoted the team to the org cap — the control widened, in the fail-open direction, on a timer, with nothing denying, so the refusal that names the deciding grain never fired precisely when a cap had stopped existing, and alert_pct / throttle_pct hung off the same row so the warning lapsed with the cap it was meant to warn about. Recorded rather than deleted: if you are reading a ledger from before 2026-09-07, this is what you are looking at
"This tenant's spend is capped", as a compositionStill not writable — but settled, not open. #3062 closed 2026-09-07 under ruling (c): the constant was renamed to stop over-claiming, rather than the composition being completed. So the composed sentence is as false as it was; what changed is that it is now a recorded naming decision instead of an unresolved concern. Do not write it

The AT LEAST wording on a 402 is not hedging. A window containing calls whose metering_confidence is not complete has a used that is a floor, not a total, so the refusal says so and counts the blind calls:

AT LEAST 812340 of 1000000 tokens used — 3 call(s) in this window carry no complete metering, so the true figure is higher

A precise number would send you looking for a billing bug. The same floor applies in the admitting direction, where it costs money rather than denying — that branch logs the incompleteness rather than staying silent (#3059).

Also enforced since #3118: all four token classes are summed (input, output, cache_creation, cache_read). Two of them were invisible to the cap before.

6.2 Rate limits — 429​

Per-identity token buckets in Redis, evaluated at step 4, before the body is read — a limiter running after a 32 MiB read would already have paid the cost it exists to bound.

ClassSustainedBurstApplies to
org200/s400every caller — this class is total
team100/s200unit-scoped credentials
agent20/s40agent-scoped credentials

Verified off the running process (rate_limits= on the wired line). There is no environment variable that changes these — deliberately, so the running value and the stated reason cannot drift. They are org × 1, × ½, × ⅒, with burst × 2, and the ratios are containment statements: a team may not out-consume its org, an agent may not out-consume its team.

The team class is deliberately oversubscribed — three teams at org/2 sum to 150% of the org rate. The per-team bucket is a blast-radius bound on one team, not a share of a fixed pie; the org bucket is what bounds the total. Partitioning instead would strand capacity whenever a team is idle and would silently change every team's limit the day a new team is created.

A caller the org class cannot bucket — no org_id — is refused, not admitted. A non-empty bucket set is not the same as a bounded caller.

Two 429 properties your client can rely on:

  • Retry-After is always present on a 429 and never on any other refusal. That is the contract, not a convention: RetryAfterRequired is declared per identity, and a header omitted at zero seconds would otherwise be indistinguishable from one never set.
  • rate_limit_unavailable (503) is not rate_limited_* (429). A limit that could not be read is not a limit that was exceeded. There is a grace window (MODEL_GATEWAY_RATE_LIMIT_GRACE_SECONDS, default 5m) during which a per-replica fallback stands in for the shared bucket; after it, step 4 fails closed.

7. Bedrock — the interim workaround​

Native AWS Bedrock is not supported and is not coming in v1 (PRD #2644 D3). This is not a gap in the routing layer; SigV4 request signing is a different auth axis and does not fit the registry's url + type + key model. Confirmed in the tree: git grep -i 'bedrock\|sigv4' returns no implementation — only ADR prose and one test fixture whose model id happens to contain the word.

The gateway speaks exactly two dialects, openai_compatible and anthropic (dialects_relayable on the wired line). So the interim path is to put something in front of Bedrock that speaks one of them.

Two options. Both are the tenant's infrastructure, not UpsQuad's.

Option A — AWS's own OpenAI-compatible endpoint​

If your account has an OpenAI-compatible chat-completions endpoint fronting Bedrock, register it directly:

  • dialect: openai_compatible
  • base_url: the endpoint's public https origin
  • credential: the endpoint's bearer key, written through SetLLMEndpointCredential

The gateway sends Authorization: Bearer <secret> and joins your request path onto base_url, so a deployment prefix in the URL works.

Option B — a LiteLLM shim you host​

Run LiteLLM configured with a Bedrock model, exposing /v1/chat/completions (or /v1/messages for the anthropic dialect), and register the shim's URL as the endpoint. LiteLLM holds the AWS credentials; UpsQuad holds a static bearer for the shim.

The constraint that catches everyone: the shim must be publicly resolvable over https. SSRF fence rule 4 refuses any host resolving into the private/loopback/ link-local/CGNAT/ULA set unconditionally, and MODEL_GATEWAY_DEV_HTTP_HOSTS — the allowlist that would exempt one — relaxes only the scheme check and is currently inert. A shim on 10.0.0.0/8, on a Kubernetes service name, or on localhost cannot be registered, and this is exactly why the devbox stack cannot relay to its own upsquad-litellm container.

infra/litellm/config.yaml is a working LiteLLM proxy config in this repo, but it is a dev-only Anthropic facade over local Ollama for the Claude-SDK worker — it is not a Bedrock configuration and it is not reachable by the Model Gateway for the reason above. Read it as a syntax reference only.

What you give up either way, and it should be said before someone finds out from a bill: metering fidelity is whatever the shim reports. If it does not return a usage block the extractor recognises, calls land as metering_confidence not-complete, and every budget number over that window becomes an AT LEAST floor (§6.1) — including the one that decides whether to admit the next call.


8. Where the spend shows up​

Tenant-facing, per org/team/agent/model — AIGatewayUsageService, read-only by construction (T16 #2918):

RPCReturns
QueryUsageUsage at org / team / agent / model / endpoint / day / session grain, clearance-gated and paginated
ExportUsageThe same as CSV
QueryLatencyPercentiles that state the population they were computed over (p99_of_measured_ms)
GetBudgetBurndownThe cap in force at one grain, and a projection that refuses or ranges whenever a cost is unknown

unknown cost is a separate bucket, never summed into the total and never dropped (OQ-3). A total that silently absorbed unpriced calls would be confidently wrong in the flattering direction on exactly the largest ones.

The five dashboards (MG-3.7) are upsquad-client work and are all delivered as of 2026-09-07: Spend + CSV/API export (client#840, merged 2026-09-05T16:24:21Z); Agent Economics, Reliability, Governance and Budget burndown (client#842, merged 2026-09-07T13:39:49Z). Both frontends carry every surface (two-frontend parity). This paragraph said the second four were "open on client#842" until that PR merged — client#842 is a merged PR, not an open issue, and MG-3.7 is complete.

Platform-operator view — Grafana, Model Gateway — Overview (deployments/observability/grafana/model-gateway-overview.json). Panels as of #3011/#3008: concurrency and in-flight upstream requests (the HPA's own input), scrape health, peak concurrency per replica, rate-limit refusals by identity class, rate-limit decisions by arm, step-4 degradation, budget refusals by deciding grain, budget decisions by arm, gate 7d failing closed, and unclassified gate-7d decisions.

Two panels exist because absence had to become visible: "gate 7d failing closed" separates a control that denied from one that could not be evaluated, and "unclassified" catches a future branch that classifies nothing rather than letting it go silent.

Denials vs allows are recorded differently, on purpose. A denial writes a llm_usage_events deny row and a hash-chained agent_audit_log hop. An allowed call writes a ledger row and no audit hop — chaining every served call would put an advisory-lock round trip on the hot path (PRD v1.2 amendment to MG-2.11).


9. What this page does not claim​

  • The hostname being live is not production traffic. The flip released E2E testing traffic under the scoped waiver on #2912; production tenant LLM traffic is still gated. #2912 itself stays open — its closure is the founder's call given the waiver scope, not a consequence of the flip.

    Do not read a gate list off this page. The register moves faster than a runbook can. On the afternoon of 2026-09-07 three things this page asserted as open closed inside eighteen minutes — client#842 merged 13:39:49Z, #2928 closed 13:55:24Z, #3062 closed 13:57:24Z. The pre-promotion checklist is a query:

    gh api "search/issues?q=repo:upsquad-ai/upsquad-core+label:prod-enablement+is:open&per_page=100" \
    --jq '.items[] | "#\(.number) \(.title)"'

    Filter it to the offering being shipped, per the Approval Protocol in CLAUDE.md. At the time of writing, #3120 is the gate whose title names this switch directly.

    And re-sweep the whole page, not the row you came for. Every state claim here carries a date for exactly this reason; the ones that rot are never the ones you are looking at:

    grep -oE '#[0-9]{3,5}' docs/runbooks/model-gateway-external-agent-onboarding.md | tr -d '#' | sort -un |
    while read -r n; do
    gh api repos/upsquad-ai/upsquad-core/issues/"$n" \
    --jq '"#\(.number) [\(.state)] closed=\(.closed_at // "-") \(.title[0:60])"'
    done

    Cross-repo references (client#…) need the same check against their repo — and check whether the number is an issue or a merged PR, because "open on client#842" was wrong in both halves.

  • The :8091 HA posture is not established. #2646 remains an open prod-enablement gate that slice 3 does not close. One devbox host edge, one process.

  • GKE ingress for the Model Gateway is undesigned. infra/edge/envoy.yaml governs the devbox host edge only; the cluster Service is ClusterIP with no Ingress, no HTTPRoute, no LoadBalancer and no ArgoCD Application, and an exposure check asserts that stays true. #2987 — deliberately not prod-enablement, because it gates no switch that exists.

    One thing that check no longer catches, recorded once so nobody rediscovers it as a surprise. scripts/check-model-gateway-exposure.py finds edge references through three EDGE_MATCHERS arms: the token model-gateway, the token model_gateway, and a digit-bounded 8091. The public hostname used to satisfy the first arm by being model-gateway.app.upsquad.ai; after the #3196 / #3198 renames it is model-gw-beta.upsquad.ai, which contains no model-gateway token. Coverage is unchanged in practice — the vhost's name: field is still model-gateway, and anything that actually reaches the gateway must carry model_gateway (the cluster) or 8091 (the port), so both other arms still fire and the census is unchanged at 4 objects. But the hostname string is no longer load-bearing for that arm. Do not prune an arm on the reasoning that "the hostname covers it".

  • The edge authn/authz/rate-limit baseline is #2985, ruled and re-scoped. Ruled option (b), gateway-enforced (founder, 2026-09-07) — that released the flip, and #2985 no longer gates it. It stays open scoped to first external-tenant onboarding. The consequence to carry: under (b) there is no Access application in front of the hostname, so the direct_response seal in §2 is the sole control on the unauthenticated ops surface. Any future edit to that vhost is a security change, not a routing change.

  • unsupported_dialect moved 502 → 501, and if you wrote a client before 2026-09-08 this is a breaking change to one branch. Founder ruling on the #2981 rider, 2026-09-08. It was served at 502, which the taxonomy maps to unavailable — "the control could not be evaluated; a retry may well succeed" — and for this refusal every control that ran said yes and no retry can succeed until a binary with a transport for the dialect ships. That is verbatim the sentence the 501 ruling declares false for the identical fact, so the code was telling clients to back off and retry a permanent condition.

    What changes on the wire: the status, and the class field, which is now absent — 501 is a declared-unclassified status, so the refusal carries no class rather than a wrong one. Nothing else. The code is unchanged, so a client keying on error.code (which is what §5 asks you to do) needs no change at all; a client branching on the status or on class == "unavailable" does.

    Why now rather than never, or later. The only public-hostname traffic today is E2E waiver traffic, so this costs its minimum; after first external-tenant onboarding it would have needed deprecation ceremony for a code nobody could have been relying on correctly.

    Its ledger outcome did not move and is still per-producer: refused pre-admission at gate 8/credential it is a deny and is in the audit chain; refused post-admission by the handler it is an allow with no hop. The class is the same at both and the disposition is not, which is not an inconsistency — the class answers "could a retry ever succeed" (no, either way) and the disposition answers "did a gate decide". Two questions.

    With that ruled, no contract disposition is contested any longer; the register that tracked them is deleted rather than left empty. One question it raises is open and tracked at #3245: the 501 ruling declines to invent a fifth class and stated it would be revisited if a second refusal in this shape appeared. Three now have. That is a vocabulary question, not a defect — the current state is a declared omission and is honest on the wire — but it has a closing window for the same reason this change had one.

  • The ledger half of #2981 is fixed, and it changes what you will find in the chain. Option B of the #2981 fork (founder ruling, 2026-09-07 — not to be confused with #2985's option (b) above, ruled the same day) moved the disposition from the refusal reason onto the producer. If you audited this gateway before that landed, two things are now different and both are corrections:

    • route_unresolved used to write governance_outcome='deny', a deciding_gate and a hash-chained audit hop for a request where no gate ever decided. It now writes a row with governance_outcome NULL — the ledger's spelling for "no gate chain evaluated this call" — and no audit hop.
    • invalid_request, request_body_too_large and unsupported_dialect are each emitted by two producers. Refused pre-admission by the pipeline they are denials and are in the chain, as before. Refused after every gate said yes they now record an allow and produce no hop, because governance permitted those calls. The deny or allow in the table above is that fact, per request.

    What this costs you as an auditor, stated plainly: the refusal reason for a non-deny row is not in deciding_gate — migration 207 permits that column only beside a deny — so it is counted on model_gateway_refusal_ledger_disposition_total{reason,disposition} and nowhere durable. That metric has a retention horizon and is not tamper-evident. The trade was made deliberately: what the chain now contains is only decisions that happened. Expect a one-off step down in governance_assessed_deny_events at the deploy and a matching step up in governance_unassessed_events; alert rules keyed on either need the note. That step is now observable rather than hypothetical: the flip released E2E testing traffic (first bullet), so these counters have a live population to move.

  • Nothing on this page was executed end-to-end through a successful governed call, for the reason in §3: the devbox has no registered external client and no registered endpoint. The flip did not change this — a live hostname that answers 401 is still a refusal path, and a 401 from the internet is not evidence of a completed governed call. Refusal shapes, the token endpoint's resource contract, the wired-controls line, the seal battery and the DNS/TLS state are all live measurements taken 2026-09-07; success shapes are read from source.