AI Gateway API
Table of Contents
-
upsquad/aigateway/v1/capability_binding.proto
-
upsquad/aigateway/v1/endpoint.proto
-
upsquad/aigateway/v1/endpoint_binding.proto
upsquad/aigateway/v1/budget.proto
BudgetSpend
BudgetSpend is the counted ledger read over the window, carrying BOTH honesty axes so neither can be inferred from the other.
Every count comes from an INDEPENDENT aggregate in SQL, so the two partition
identities (priced + unknown_cost == calls, metered + unmetered == calls)
are real checks rather than arithmetic tautologies. The service verifies both
before answering and REFUSES the call if either fails — a wrong spend figure
that renders is worse than one that errors.
| Field | Type | Label | Description |
|---|---|---|---|
| calls | int64 | calls is the number of ledger rows in the window at this grain. | |
| priced_events | int64 | priced_events is the number of calls with a known cost. | |
| unknown_cost_events | int64 | unknown_cost_events is the number of calls whose cost is UNKNOWN (llm_usage_effective_cost.cost_unknown). It is a separate counter that no arithmetic on any money field can absorb. | |
| metered_events | int64 | metered_events is the number of calls whose token counts are TOTALS (metering_confidence = 'complete'). | |
| unmetered_events | int64 | unmetered_events counts calls whose metering_confidence is anything else — NULL, 'partial' or 'blind'. When it is > 0, tokens_used is a LOWER BOUND and so is every percentage computed from it. | |
| spend_usd | string | optional | spend_usd is the summed cost over PRICED events only, as a decimal string. |
ABSENT — never "0.00000000" — when the window holds no priced event. That distinction is the whole of OQ-3 and it is the only input on which an honest sum and a COALESCE-ing one disagree. |
| tokens_used | int64 | | tokens_used is input + output tokens over the window: the quantity the org and team TOKEN caps are measured against. A LOWER BOUND whenever unmetered_events > 0. |
BurndownProjection
BurndownProjection is the exhaustion projection, shaped so that refusing is representable.
READ status FIRST. The money fields are meaningless without it: on REFUSED
they are empty strings, and an empty string parsed as a number is 0.00 —
exactly the fabricated zero this message exists to prevent.
THE TWO MONEY FIELDS DO NOT HAVE EQUAL STANDING, AND THEIR NAMES SAY SO.
lower_bound_usd is proven. modelled_ceiling_usd is an estimate biased
toward under-reporting burn — the fail-open direction — and it is named for
what it is rather than paired symmetrically with the bound. Read its comment
before writing any rule that compares it to a cap.
| Field | Type | Label | Description |
|---|---|---|---|
| status | ProjectionStatus | status is the epistemic status of this projection. See ProjectionStatus. | |
| lower_bound_usd | string | lower_bound_usd is a PROVEN floor on the window's true spend: the sum over priced calls. True spend is greater than or equal to it, because an unpriced call's cost is unknown but non-negative. It is the only number in this message that earns the word "bound". |
Decimal string at NUMERIC(14,8) scale. EMPTY on REFUSED. | | modelled_ceiling_usd | string | | modelled_ceiling_usd IS NOT AN UPPER BOUND AND IS DELIBERATELY NOT NAMED LIKE ONE. It is an ESTIMATE, and it is BIASED LOW — toward under-reporting burn, which is the fail-open direction.
WHY THE NAME IS THE CONTROL HERE
This field was called upper_bound_usd, paired symmetrically with lower_bound_usd. Symmetric names assert symmetric epistemic status, and everything that corrected the impression — the field comment, ceiling_basis, the RANGE_UNKNOWN_COST status — sits where a consumer may never look. THE FIELD NAME IS THE ONE THING EVERY CONSUMER READS: the alert rule, the CSV header, the dashboard binding, the SQL over the export. The Go struct and the SQL column already called it modelled_ceiling; the honest name was being lost exactly at the wire boundary.
THE BIAS IS DIRECTIONAL, AND IT DEFEATS THE CONTROL
A call is unpriced BECAUSE ITS MODEL HAS NO PRICING CARD, while the witness this figure is derived from is drawn from the models that DO. Nothing makes an unpriced model cheaper than the most expensive priced one, and the unpriced population — new, unusual, frontier — skews expensive. So the estimate errs toward UNDER-estimating, in exactly the case that matters:
1 priced embedding at $0.01, 1 unpriced call that really cost $8.00 lower_bound_usd 0.01000000 true, proven modelled_ceiling_usd 0.02000000 floor + 1 x the priced witness true spend 8.01 against a $1.00 cap
A consumer comparing this figure to a cap would conclude "worst case, within cap" while 8x over. That is the same class as valuing unknowns at zero — which is merely the degenerate member of this family — one layer up.
THE RULE THAT FOLLOWS
THIS FIGURE MAY BE BELOW TRUE SPEND AND MUST NOT BE USED TO CONCLUDE THAT A CAP IS RESPECTED. Use it to show how wide the unknown is, never to clear a budget. unknown_cost_events > 0 is the signal that the true number is unbounded above; only lower_bound_usd supports a conclusion, and only the conclusion "at least this much has been spent".
Decimal string at NUMERIC(14,8) scale. EMPTY on REFUSED. Equal to lower_bound_usd on EXACT, where there is nothing to model and nothing is therefore biased. |
| unknown_cost_events | int64 | | unknown_cost_events is carried HERE as well as on BudgetSpend, deliberately.
It travels WITH the bounds so a client that renders only the projection cannot lose it, and so no arithmetic on the bounds can absorb it. Two copies of a count are ordinarily a defect; here the duplication is the point, and the service asserts the two agree before answering. |
| ceiling_basis | string | | ceiling_basis names how modelled_ceiling_usd was derived, including the witness value it was derived from so a reader can check the arithmetic, AND the DIRECTION of its bias.
The direction is stated in the string itself and not only in this comment, because a figure travelling to a log, an alert body or a support ticket arrives without its schema. A reader who sees only the two strings must still be told the ceiling can sit below the truth.
Non-empty EXACTLY on RANGE_UNKNOWN_COST. Empty on EXACT (nothing was modelled) and on REFUSED (nothing was produced). | | refusal_reason | string | | refusal_reason names why no projection exists.
Non-empty EXACTLY on REFUSED. It is prose for an operator, not a code: a client MUST branch on status, never on this string. |
GetBudgetBurndownRequest
GetBudgetBurndownRequest asks for one grain's burndown.
THERE IS NO ORG FIELD AND THERE MUST NEVER BE ONE. The org comes from the authenticated caller. A body-nameable org is a cross-tenant read wearing the costume of a filter parameter, and the absence of the field is what makes that unconstructable rather than merely refused.
There is no window field either. The window is the BUDGET PERIOD IN FORCE at
the requested grain — a burndown against an arbitrary caller-chosen window is
not a burndown, and a caller who wants an arbitrary window has
AIGatewayUsageService.QueryUsage.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain is the budget grain. Required; UNSPECIFIED and any unmapped number are refused with InvalidArgument. | |
| team_id | string | team_id selects the team when grain is BUDGET_GRAIN_TEAM. |
REQUIRED at team grain and REFUSED at every other grain. A team_id beside BUDGET_GRAIN_ORG is a request that means two different things, and answering the org total under a team label is the "answers a different question" defect the grain enum is closed to prevent.
It is matched inside the caller's org, so a team id belonging to another org selects nothing rather than leaking. |
GetBudgetBurndownResponse
GetBudgetBurndownResponse is one grain's cap, its honest burn, and a projection that can refuse.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain echoes the grain that was answered, so a response cannot be attributed to a different question than it answers. | |
| bucket_key | string | bucket_key identifies WHAT was answered: the org id at org grain, the team id at team grain. Echoed from what the server resolved, never from the request. | |
| window_start | google.protobuf.Timestamp | window_start is the inclusive start of the budget period in force. |
Present whenever a window could be determined at all — including on BUDGET_CAP_STATE_LAPSED, where it reports the STORED window so an operator can see how stale it is. Absent on BUDGET_CAP_STATE_UNREADABLE, which is precisely the state of having no derivable window. |
| window_end | google.protobuf.Timestamp | | window_end is the EXCLUSIVE end of the budget period in force. Same presence rules as window_start. |
| cap_state | BudgetCapState | | cap_state says whether a cap is in force, and when it is not, why. A client MUST branch on it before reading the limits: UNCONFIGURED and LAPSED both mean "nothing is capping right now" and they are not the same fact. |
| token_limit | int64 | optional | token_limit is the cap in tokens — feature_gates.limit_value at org grain, team_budgets.budget_tokens at team grain.
ABSENT when no limit is configured. Absent at org grain is also the enterprise tier's NULL, which genuinely means unlimited; cap_state disambiguates. |
| usd_limit | string | optional | usd_limit is team_budgets.budget_usd — the OPTIONAL dollar cap, absent at org grain (there is no dollar gate in feature_gates) and absent for a team that tracks tokens only.
Decimal string. Where this is present, the fabricated-zero defect is not a reporting error but a cap that silently does not fire. |
| spend | BudgetSpend | | spend is the counted ledger read over the window. Always present, including the all-zero window: an omitted spend is indistinguishable from a client that failed to read it, and "nothing was spent this period" is an answer a dashboard must be able to render confidently. |
| projection | BurndownProjection | | projection is the exhaustion projection. Always present, and always carrying a status — a response with no projection would put the client back in the position of inferring one. |
BudgetCapState
BudgetCapState says WHETHER a cap is in force at the requested grain, and when it is not, WHY.
The distinction between UNCONFIGURED and LAPSED is the #2928 read constraint made expressible: both render as "no limit" if collapsed, and only one of them means the tenant was never given one.
| Name | Number | Description |
|---|---|---|
| BUDGET_CAP_STATE_UNSPECIFIED | 0 | BUDGET_CAP_STATE_UNSPECIFIED is never sent. It exists because proto3 requires a zero value, and a response carrying it is a server fault. |
| BUDGET_CAP_STATE_IN_FORCE | 1 | BUDGET_CAP_STATE_IN_FORCE means a cap row exists and the current instant falls inside its window. This is the only state in which the enforcement gate would apply this cap. |
| BUDGET_CAP_STATE_UNCONFIGURED | 2 | BUDGET_CAP_STATE_UNCONFIGURED means the grain was asked and there is no cap there. At org grain this is also the enterprise tier, whose feature_gates.limit_value is NULL meaning genuinely unlimited. |
| BUDGET_CAP_STATE_LAPSED | 3 | BUDGET_CAP_STATE_LAPSED means a cap row EXISTS and its stored [period_start, period_end) does not contain the current instant, so no cap is in force right now. |
This is NOT "unlimited" and must not be rendered as headroom. The window is still reported so an operator can see how far out of date it is. #2928's product ruling makes these windows recur monthly on the org billing anchor; that ships with the write surface, and until it does this state is the honest report of what the enforcement gate actually does today. |
| BUDGET_CAP_STATE_UNREADABLE | 4 | BUDGET_CAP_STATE_UNREADABLE means a LIMIT exists but NO WINDOW can be derived for it — subscriptions.billing_anchor_at is NULL, so billing_period_bounds (a STRICT function) returns no bounds.
UNKNOWN, NEVER UNLIMITED (#2861). Composing a real limit against a sum over an unbounded or empty predicate reads as full headroom and disables the cap silently. A burndown in this state REFUSES its projection. |
BudgetGrain
BudgetGrain is the grain a burndown is asked at.
It is a SEPARATE, CLOSED enum and deliberately not UsageGrain. A budget
exists at org, team and agent grain and nowhere else; UsageGrain is a GROUP
BY dimension that is growing members (ENDPOINT, DAY, SESSION, CAPABILITY),
and reusing it would make GetBudgetBurndown(grain: DAY) representable — a
request that cannot mean anything, permanently one refusal away from being
answered by a default.
An unmapped NUMBER is refused with InvalidArgument and NEVER falls through to
BUDGET_GRAIN_ORG. proto3 enums are OPEN, so a newer or hand-rolled client can
put an undeclared number on the wire; the compatibility contract MG-3.7 T2
(#3024) pinned for UsageGrain binds identically here.
| Name | Number | Description |
|---|---|---|
| BUDGET_GRAIN_UNSPECIFIED | 0 | BUDGET_GRAIN_UNSPECIFIED is REFUSED with InvalidArgument. It is not a default: answering the org's burndown for a caller who asked for a team's is indistinguishable from a correct answer at the call site. |
| BUDGET_GRAIN_ORG | 1 | BUDGET_GRAIN_ORG is the whole organisation, capped by feature_gates.limit_value for llm_tokens_monthly at the org's plan tier, over the period derived from subscriptions.billing_anchor_at (migration 223). It takes no team_id. |
| BUDGET_GRAIN_TEAM | 2 | BUDGET_GRAIN_TEAM is one team, capped by that team's team_budgets row (migration 025). It REQUIRES team_id. |
| BUDGET_GRAIN_AGENT | 3 | BUDGET_GRAIN_AGENT is the single agent's cap — llm_agent_budgets (migration 212). It is DECLARED AND DELIBERATELY NOT SERVED BY GetBudgetBurndown, and it is declared precisely so that fact is legible on the wire rather than being an absence a caller has to infer. |
THE REDIRECT THAT USED TO BE HERE WAS WRONG (#3064)
This comment used to send a caller to upsquad.agent.v1.AgentService.GetBudgetStatus as "the agent grain's existing surface". IT IS A DIFFERENT PRODUCT. That RPC reads agent_budgets (migration 057) — daily_token_cap / monthly_token_cap / daily_cost_cap / monthly_cost_cap, reset daily — while gate 7d enforces llm_agent_budgets.budget_tokens over the org's billing period. An agent refused with 402 budget_exceeded_agent was being sent to a surface reporting a cap the gateway does not enforce, over a window it does not use.
Neither surface was wrong on its own terms; the REDIRECT BETWEEN THEM was, and it read as authoritative because it was written as a deliberate choice. It is removed rather than softened.
upsquad.aigateway.v1.AIGatewayGrainBudgetService.GetGrainBudget is the surface for THIS cap (#2928). Asking for the agent grain here is refused with InvalidArgument naming that RPC — a strictly more useful answer than "unknown grain", and now a correct one.
A burndown at agent grain — cap, counted spend, refusing projection — is still not served by anything, and that absence is recorded rather than papered over: it needs the ledger read GetGrainBudget deliberately does not do. #3064 tracks the choice between serving it here and stating plainly that the two are different products; this comment does the latter now. |
ProjectionStatus
ProjectionStatus is the epistemic status of the burndown projection.
It is the field that makes refusal representable. See this file's header for why a scalar projection is a control that fails open.
| Name | Number | Description |
|---|---|---|
| PROJECTION_STATUS_UNSPECIFIED | 0 | PROJECTION_STATUS_UNSPECIFIED is never sent. A response carrying it is a server fault, and a client MUST treat it as REFUSED rather than as EXACT. |
| PROJECTION_STATUS_EXACT | 1 | PROJECTION_STATUS_EXACT means every call in the window is priced, so spend to date is a TOTAL and not a floor. lower_bound_usd and modelled_ceiling_usd are equal — nothing was modelled — and unknown_cost_events is 0. This is the ONLY status under which a consumer may compare a spend figure to a cap and act on the answer. |
IT IS UNREACHABLE WHENEVER unknown_cost_events > 0. That is the invariant this whole file exists to hold. |
| PROJECTION_STATUS_RANGE_UNKNOWN_COST | 2 | PROJECTION_STATUS_RANGE_UNKNOWN_COST means some calls are unpriced and an estimate could be derived, so the answer is a RANGE. lower_bound_usd is strictly the priced floor and is PROVEN; modelled_ceiling_usd is that floor plus the unpriced calls valued by ceiling_basis and is an ESTIMATE BIASED LOW. The range widens as the unknowns grow.
TRUE SPEND IS NOT BOUNDED ABOVE UNDER THIS STATUS. See modelled_ceiling_usd for why the ceiling may sit below true spend, and why no consumer may use it to conclude a cap is respected. |
| PROJECTION_STATUS_REFUSED | 3 | PROJECTION_STATUS_REFUSED means no honest projection exists. Both money fields are EMPTY STRINGS — not "0.00000000" — and refusal_reason names the cause.
Reached when the window holds unpriced calls and NO priced call to derive a ceiling from (in particular, a window in which EVERY call is unpriced), and when the cap itself is unreadable. |
AIGatewayBudgetService
AIGatewayBudgetService is the tenant-facing READ surface over the org and team budget grains.
The org is taken from the AUTHENTICATED CALLER and is NOT NAMEABLE BY THE
REQUEST BODY — GetBudgetBurndownRequest declares no org field at all, which is
a structural guarantee rather than a validation rule, and
TestBudgetBurndown_RequestDeclaresNoOrgField derives it from the message
descriptor so it cannot be reintroduced by a field addition.
READ-ONLY BY CONSTRUCTION. No RPC here writes anything — see the #2928 section of this file's header for why that is an acceptance criterion and not a description.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| GetBudgetBurndown | GetBudgetBurndownRequest | GetBudgetBurndownResponse | GetBudgetBurndown returns the cap in force at one grain, the honest spend and token burn against it, and a projection that REFUSES OR RANGES rather than fabricating a total when any call in the window is unpriced. |
Requires L3, the same floor as AIGatewayUsageService.QueryUsage: this surface reveals the org's own spend and its own caps, and a higher bar would put a team's budget out of reach of the person who runs that team. |
upsquad/aigateway/v1/capability_binding.proto
CapabilityBinding
CapabilityBinding is one configured capability.
There is no unit_id field — see the file header. There is no credential
field and there never will be one.
| Field | Type | Label | Description |
|---|---|---|---|
| id | string | id is the binding row's UUID. It is returned for correlation with audit records and is NOT accepted on any request: capability is the key. | |
| capability | string | capability is one of ADR-0035's five STABLE STRINGS: "embedding", "compaction", "extraction", "moderation", "rag_generation". | |
| endpoint_id | string | endpoint_id is the registered LLM endpoint this capability dials. | |
| endpoint_name | string | endpoint_name is that endpoint's operator-chosen label, resolved at read time. Carried so a card can render without a second round trip, and so an audit reader is not left holding a UUID for an endpoint that has since been deregistered. | |
| endpoint_status | string | endpoint_status is the endpoint's registry status — "pending" | |
| model | string | model is the model this capability calls. Always non-empty: migration 239's lcb_model_present_check refuses "", because "" would mean "whatever the endpoint defaults to" — the process-wide default this table exists to replace, wearing a configured-looking value. |
For capability "embedding" this is MG-CTX.22's stamp. Changing it on an org with a non-empty corpus is a RE-EMBEDDING decision: the corpus keeps its rows under the old model and recall filters the ANN on model_version, so ordinary ingests then refuse (5B3 #3156's embeddingstamp guard) until rag-reembed has run. This service does not refuse the configuration change — refusing it would make re-embedding impossible, since the new model must be configured BEFORE the re-embed can be run. |
| enabled | bool | | enabled is false when the operator has switched this capability off. Gate 7a treats false exactly as "no binding". |
| configured_by | string | | configured_by is the member id who last wrote this row. |
| created_at | google.protobuf.Timestamp | | created_at is when the capability was first configured. |
| updated_at | google.protobuf.Timestamp | | updated_at is when it was last written — by an edit or by an enable/disable. |
CapabilityCard
CapabilityCard is one of MG-CTX.6's five cards: a capability, and its binding if it has one.
The two-field shape is deliberate. A bare optional CapabilityBinding would
make "unconfigured" and "the client failed to parse it" the same value on the
wire; configured states the fact positively so an empty card is a claim
rather than an absence.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability is the member of the closed set this card is for. Present on every card, configured or not — it is the card's identity. | |
| configured | bool | configured is true exactly when binding is set. | |
| binding | CapabilityBinding | binding is the configuration, unset when configured is false. |
CapabilityImpact
CapabilityImpact is one capability that an action on an endpoint would affect.
enabled is carried because a refusal or a warning that cannot distinguish
"this endpoint serves your embedding capability" from "this endpoint is your
embedding capability, currently switched off" tells the operator to do two
different things in the two cases.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability is one of the five STABLE STRINGS. | |
| enabled | bool | enabled is the binding's current state. A disabled binding is still an impact: it is configuration the operator means to switch back on. |
CreateCapabilityBindingRequest
CreateCapabilityBindingRequest configures a capability for the first time.
There is no org_id field. There is no unit_id field. There is no enabled
field — a create lands ENABLED, because configuring a capability into an
off state is a two-step the product has no use for, and an enabled=false
default would be a configured capability that silently does nothing.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability is one of the five STABLE STRINGS. An unrecognised value is InvalidArgument naming the valid set — never a default and never a silently-dropped row. | |
| endpoint_id | string | endpoint_id is a registered, non-deregistered endpoint in the caller's org. An endpoint belonging to another org is NotFound, and it is NotFound because the INSERT sources it from an org-guarded SELECT rather than because a handler checked — a guard in a handler is a guard the next handler forgets (#2828's ruling one table over). | |
| model | string | model is the model this capability will call. Required and non-empty after trimming. |
CreateCapabilityBindingResponse
CreateCapabilityBindingResponse returns the new row.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | CapabilityBinding | binding is the configured capability. |
DeleteCapabilityBindingRequest
DeleteCapabilityBindingRequest unconfigures one capability.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability names which of the five. The key, as everywhere on this service: there is no binding id on the wire, so this request cannot name another tenant's row even in principle. | |
| reason | string | reason is an optional operator note recorded in the audit detail. It is the field that answers "why was this taken away" for the one action on this service whose subject no longer exists afterwards. |
DeleteCapabilityBindingResponse
DeleteCapabilityBindingResponse returns what was removed.
| Field | Type | Label | Description |
|---|---|---|---|
| deleted_binding | CapabilityBinding | deleted_binding is the configuration AS IT WAS at the moment of deletion. |
Returned rather than an empty ack because it is the operator's only remaining copy of what they just unconfigured — and because a client that renders it can show "you removed: <endpoint> / <model>" instead of a bare success. |
DescribeEndpointCapabilityImpactRequest
DescribeEndpointCapabilityImpactRequest names the endpoint under consideration.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id is the endpoint the operator is about to disable or deregister. |
DescribeEndpointCapabilityImpactResponse
DescribeEndpointCapabilityImpactResponse enumerates the blast radius.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id echoes the request. | |
| endpoint_name | string | endpoint_name is the endpoint's label, so the dialog can name it. | |
| impacts | CapabilityImpact | repeated | impacts is every capability binding in this org that references the endpoint, in capability order — up to five. EMPTY means nothing would be taken down, which is a positive answer and not a missing one. |
ListCapabilityBindingsRequest
ListCapabilityBindingsRequest takes nothing.
The org comes from the auth context. There is no filter, no page size and no cursor: the answer is a five-member closed set and every field that could shrink it is a field that could hide a card.
ListCapabilityBindingsResponse
ListCapabilityBindingsResponse carries exactly five cards.
| Field | Type | Label | Description |
|---|---|---|---|
| cards | CapabilityCard | repeated | cards is one entry per member of ADR-0035's closed set, in card order — embedding, compaction, extraction, moderation, rag_generation. Length is always the size of the declared set. |
SetCapabilityBindingEnabledRequest
SetCapabilityBindingEnabledRequest turns one capability off or back on.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability names which of the five. The key, as everywhere on this service. | |
| enabled | bool | enabled is the target state. | |
| reason | string | reason is an optional operator note recorded in the audit detail. |
SetCapabilityBindingEnabledResponse
SetCapabilityBindingEnabledResponse returns the row after the write.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | CapabilityBinding | binding is the configuration at its new state. | |
| changed | bool | changed is false when the request named the state the binding was already in. The call is ACCEPTED and AUDITED either way — pressing a kill-switch is worth recording even when it was already down — but changed reports honestly that nothing moved. This mirrors SetLLMEndpointStatus exactly. |
UpdateCapabilityBindingRequest
UpdateCapabilityBindingRequest repoints a configured capability.
| Field | Type | Label | Description |
|---|---|---|---|
| capability | string | capability names which of the five to update. It is the key, not an editable field: moving a configuration from one capability to another is a create and a delete, not an update. | |
| endpoint_id | string | endpoint_id is the new endpoint. Applied only when endpoint_id_set. | |
| endpoint_id_set | bool | endpoint_id_set carries PRESENCE. Without it, "" is indistinguishable from "not supplied", and a partial edit would silently clear the field it did not mention. | |
| model | string | model is the new model. Applied only when model_set. |
For "embedding" on an org with a corpus, this is a re-embedding decision — see CapabilityBinding.model. The change is ACCEPTED and AUDITED; the consequence is enforced at the write path by 5B3's guard, and surfaced by 5E4's readiness report and 5B4's card. | | model_set | bool | | model_set carries PRESENCE, for the same reason endpoint_id_set does. |
UpdateCapabilityBindingResponse
UpdateCapabilityBindingResponse returns the row after the edit.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | CapabilityBinding | binding is the updated configuration. | |
| changed_fields | string | repeated | changed_fields lists the columns the update actually wrote. An update whose values all matched is accepted and audited with an EMPTY list — the call happened, and reporting that nothing moved is more honest than reporting a change that did not. |
AIGatewayCapabilityBindingService
AIGatewayCapabilityBindingService is the org's central context configuration: which endpoint and model each of the five platform capabilities uses.
Every RPC is scoped to the caller's org, taken from the auth context and
NEVER from a request body. Reads and writes run inside the request's
RLS-bound scope transaction, so migration 239's
tenant_isolation_llm_capability_bindings policy is the isolation mechanism —
AND every statement additionally carries an explicit org_id = $1 predicate,
because RLS is skipped entirely on a BYPASSRLS connection, the class operator
CLIs and tools actually use (#3228/#3229, census #3230).
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| ListCapabilityBindings | ListCapabilityBindingsRequest | ListCapabilityBindingsResponse | ListCapabilityBindings returns MG-CTX.6's five capability cards. |
IT ALWAYS RETURNS EXACTLY FIVE, one per member of ADR-0035's closed set, in card order, whether configured or not. An unconfigured capability is a card with configured = false and no binding — not an absent card.
That is a requirement and not a convenience. A response that omitted unconfigured capabilities would let the tab render four cards on an org whose fifth is silently unconfigured, which is the "absence scored as a pass" family this milestone has paid for most: the operator sees a complete-looking screen and the missing capability is discovered when it fails. There is no pagination for the same reason — a five-member closed set has no page two, and a page two is a way for a card to go missing.
Requires L3. | | DescribeEndpointCapabilityImpact | DescribeEndpointCapabilityImpactRequest | DescribeEndpointCapabilityImpactResponse | DescribeEndpointCapabilityImpact answers, BEFORE the operator acts, what disabling or deregistering an endpoint would take down.
MG-CTX.9's second constraint: sharing is permitted, so one endpoint may serve up to FIVE capabilities, and a single kill-switch press can take embedding, compaction, extraction, moderation and RAG chat generation down together. "Are you sure?" without enumerating that is not a control.
This is the read the confirm dialog calls. SetLLMEndpointStatus computes the SAME list server-side and returns it on the response, so a client that skipped the preview still gets the answer and the audit record still names the blast radius. One computation, two surfaces — not two enumerations that could disagree.
Requires L3. | | CreateCapabilityBinding | CreateCapabilityBindingRequest | CreateCapabilityBindingResponse | CreateCapabilityBinding configures a capability for the first time.
It is NOT an upsert: an already-configured capability is AlreadyExists, and the operator is told to use UpdateCapabilityBinding. Silently overwriting an existing central configuration on a create is how one operator's first-time setup redirects another's live embedding traffic.
Requires L5. |
| UpdateCapabilityBinding | UpdateCapabilityBindingRequest | UpdateCapabilityBindingResponse | UpdateCapabilityBinding repoints a configured capability at a different endpoint and/or model. Fields are optional-by-presence via *_set companion flags, matching UpdateLLMEndpoint's shape.
It does NOT write enabled — that is SetCapabilityBindingEnabled's, so "who turned this capability off" is answerable by one audit predicate rather than by reading the changed-field list of every update.
Requires L5. | | SetCapabilityBindingEnabled | SetCapabilityBindingEnabledRequest | SetCapabilityBindingEnabledResponse | SetCapabilityBindingEnabled turns one capability's binding off or back on.
A DISABLED BINDING IS NOT A DELETED ONE. Gate 7a treats disabled exactly as missing — ErrCapabilityUnconfigured, 403 — but the configuration survives, so re-enabling does not cost the operator a re-entry, and DeregisterLLMEndpoint still refuses underneath it (a disabled binding is configuration the operator means to switch back on).
Requires L5. | | DeleteCapabilityBinding | DeleteCapabilityBindingRequest | DeleteCapabilityBindingResponse | DeleteCapabilityBinding UNCONFIGURES a capability entirely.
DISABLE IS THE ORDINARY CONTROL AND THIS IS NOT A SECOND SPELLING OF IT. Disable keeps the configuration so it can be switched back on without re-entry; delete removes the row, and the capability returns to configured = false on its card and bound = false in the readiness report.
It exists because 5B1's DeregisterLLMEndpoint refuses while ANY binding references the endpoint — enabled or disabled, deliberately. Without a delete, binding an endpoint once would make that endpoint permanently underegisterable, with the refusal naming a row the operator has no surface to remove. A control whose only escape is a DBA is not a control.
The history is NOT lost: the audit row records what was unconfigured — endpoint, model and enabled-state at the moment of deletion — which is strictly more than the row itself held.
Requires L5. |
upsquad/aigateway/v1/credential_verification.proto
VerifyEndpointCredentialRequest
VerifyEndpointCredentialRequest names ONE (endpoint, grain) pair and nothing else.
THREE FIELDS, AND THE ABSENCES ARE THE CONTRACT. No org_id (it comes from the token), no base_url / dialect / auth_mode (they come from the row), no secret, no body, no method, no path, no timeout. Every one of those, if accepted here, would widen this from "re-check a credential the org already stored against an endpoint the org already registered" into something the census was right to refuse.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id is llm_endpoints.id. The ROW is read under the token's org (V3), so an endpoint in another org answers NotFound — identically to one that does not exist. |
It MUST equal the endpoint_id the token is bound to (V1). A token is minted for one pair, so a captured token can at most re-probe that pair. | | grain | string | | grain is the vault slot, in credentialscope's vocabulary: "_default" (org) | "<unit uuid>" (team) | "agent<uuid>" (agent).
It is the vault's team_id column verbatim and is never re-spelled or re-derived server-side — it arrives already resolved by the caller that wrote the credential. It MUST equal the grain the token is bound to (V1). |
| trigger | Trigger | | trigger is why. Audit detail only; TRIGGER_UNSPECIFIED is refused. |
VerifyEndpointCredentialResponse
VerifyEndpointCredentialResponse is what one probe found and what was written.
IT REPORTS THE WRITES SEPARATELY FROM THE FINDINGS (recorded,
models_recorded), because a probe that learned the truth and failed to
persist it is a different operational state from one that never ran — and the
badge an operator is looking at is the persisted one.
| Field | Type | Label | Description |
|---|---|---|---|
| status | string | status is the recorded vocabulary value: "verified" |
2xx IS THE ONLY SUCCESS BAND. 3xx is deliberately NOT folded into it: the production transport refuses to follow redirects (it would send the tenant's credential to a Location the upstream chose), so a 3xx on a real request does not complete — badging it verified would bless a credential that cannot serve a turn.
IT IS NEVER "stale". stale means "not verified against the CURRENT base_url", which is a statement about the endpoint row changing, not about anything a probe observed. Only the context-engine writes it, in the same transaction as the URL change. |
| models_discovered | bool | | models_discovered distinguishes "we did not learn the catalogue" from "the provider serves nothing".
false ⇒ nil, and the operator's typed list must stand: a non-2xx, an envelope neither dialect documents, or a truncated body all land here. true + empty discovered_models ⇒ the provider answered 200 with {"data":[]}, i.e. it really serves nothing — a reason to worry rather than a reason to shrug. Collapsing the two would erase exactly that distinction, and the erasure is silent in the direction that wipes a good catalogue on a transient 503. |
| discovered_models | string | repeated | discovered_models is data[].id and NOTHING ELSE.
Both dialects document the same envelope, so one parser serves both; Anthropic adds type/display_name/created_at and OpenAI adds object/created/owned_by, and none of it is read. A picker needs the id, and every additional field kept is a field that could carry something this platform should not store. |
| probed_at | google.protobuf.Timestamp | | probed_at is when the probe was taken, UTC. Zero when no probe ran. |
| recorded | bool | | recorded reports whether the verification outcome reached provider_keys.last_verify_status / last_verified_at (migration 259).
A false here with a non-empty status means the probe ran and the badge did not move — the RPC still returns OK, because the credential write this rode along with has already succeeded and an advisory badge is not worth failing it. It is logged server-side, because a badge silently reading "never probed" is the #3581 shape wearing new clothes. |
| models_recorded | bool | | models_recorded reports whether the catalogue reached llm_endpoints.discovered_models / discovered_at (migration 260).
THE CATALOGUE WRITE RULE, which settles the last-writer-wins race the #3594 review found: the _default grain ALWAYS writes; any other grain writes only when discovered_at IS NULL. Three grains probing one endpoint would otherwise leave the picker showing whichever grain finished last, and the org grain is the one whose catalogue an operator means. |
| notices | string | repeated | notices are operator-facing sentences.
HOST ONLY, NEVER A PATH OR A QUERY, AND NEVER ANY SECRET MATERIAL. A base_url carrying an api-version query must not be able to leak into a message, so every rendering goes through a host-only helper.
A refusal notice always names BOTH the key and the base_url, in the same breath, because the probe genuinely cannot tell them apart — #3581's key was CORRECT and the url was wrong, and a bare "401" sends an operator to re-paste a key for fourteen days. |
Trigger
Trigger names WHY a verification was attempted. AUDIT DETAIL ONLY — it selects no behaviour, and V1–V8 run identically whatever it says.
It exists because "who re-probed this endpoint forty times" and "what was happening when this key first failed" are the two questions the audit row gets asked, and neither is answerable from a timestamp alone.
| Name | Number | Description |
|---|---|---|
| TRIGGER_UNSPECIFIED | 0 | TRIGGER_UNSPECIFIED is refused rather than defaulted. A caller that did not say why it is probing is a caller whose audit row would claim a reason it does not have, and MANUAL is a specific claim about a human acting. |
| TRIGGER_CREDENTIAL_SET | 1 | TRIGGER_CREDENTIAL_SET — SetLLMEndpointCredential stored a credential at this grain for the first time, or replaced one. |
| TRIGGER_CREDENTIAL_ROTATED | 2 | TRIGGER_CREDENTIAL_ROTATED — the same grain's secret was rotated. |
| TRIGGER_ENDPOINT_CHANGED | 3 | TRIGGER_ENDPOINT_CHANGED — the endpoint's base_url or dialect moved, so every recorded verification was marked stale and this is the post-commit sweep re-probing one grain against the NEW url. |
| TRIGGER_MANUAL | 4 | TRIGGER_MANUAL — an operator asked, through AIGatewayEndpointCredentialService.VerifyLLMEndpointCredential. #3581 recorded that no "test connection" action existed; #3600 wants a picker "refresh". Both are that one RPC, and it is what makes this path testable without staging a URL change. |
ModelGatewayCredentialVerificationService
ModelGatewayCredentialVerificationService verifies one stored endpoint credential by using it, once, against the endpoint it belongs to.
Served by cmd/model-gateway on its internal listener. Called by cmd/context-engine from the credential write paths. There is deliberately ONE RPC: every caller wants the same thing, and a second RPC would be a second surface to gate.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| VerifyEndpointCredential | VerifyEndpointCredentialRequest | VerifyEndpointCredentialResponse | VerifyEndpointCredential performs exactly one authenticated GET {base_url}/{version}/models with the credential stored at grain for endpoint_id, records the outcome and the discovered catalogue, and writes one chained audit row. |
── A PROVIDER REFUSAL IS NOT AN ERROR ────────────────────────────────────
status = "401" with code OK is the normal, expected answer for a bad key. That is the whole point of the RPC: the finding is the RESULT, not a failure of the call. Error codes are reserved for the V-chain refusing, or for a dependency being down:
Unauthenticated V1 — signature, issuer, expiry, or aud is not EXACTLY the credential-verification audience. PermissionDenied V2 — clearance below WriteMinClearance; OR the request's (endpoint_id, grain) differs from the pair the token is bound to. NotFound V3 — no such endpoint IN THE CALLER'S ORG. An endpoint belonging to another org is deliberately indistinguishable from one that does not exist. ResourceExhausted V4 — 30/endpoint/minute or 120/org/minute exceeded. Unavailable Redis, Postgres or the vault is unreachable. The gate FAILS CLOSED: no probe runs. "The rate limiter cannot answer" is not permission to skip the rate limit. FailedPrecondition V5 — unknown auth_mode, a dialect this build has no transport for, or no credential stored at grain.
NOTHING HERE EVER RETURNS SECRET MATERIAL, and no field, notice, log line or audit row carries the response body, the URL path or the URL query. |
upsquad/aigateway/v1/endpoint.proto
CredentialSummary
CredentialSummary is the ONLY projection of an endpoint credential that ever crosses the database boundary (LLD §8.4, normative).
There is no Get. There is no read-back. Ever. The mask is computed inside Postgres, so the plaintext is never materialised in Go process memory and never reaches the client. The frontend must never hold a secret in client state — the write is fire-and-forget and the response carries only the hint.
This message is declared here rather than in endpoint_credential.proto because LLMEndpoint embeds it and this file lands first; T10's credential service reuses this exact type.
| Field | Type | Label | Description |
|---|---|---|---|
| present | bool | present is derived from row existence via a non-decrypting vault probe. | |
| masked_hint | string | masked_hint is "••••" followed by the secret's last 4 characters, or bare "••••" when the secret is 4 characters or shorter, or "" when absent. | |
| last_rotated | google.protobuf.Timestamp | last_rotated is when the secret was most recently written. Unset when absent. | |
| last_verified_at | google.protobuf.Timestamp | last_verified_at is when the gateway last PROBED this credential grain against the endpoint's stored base_url (#3581). Unset means never probed — which is the state of every credential written before #3581 and of any deployment whose probe could not run. |
UNSET IS NOT "BAD". A client must render it as "not verified", never as a failure: a credential that has never been probed is exactly as likely to work as one probed successfully a week ago. | | last_verify_status | string | | last_verify_status is the outcome of that probe. FOUR SHAPES, and a client must handle an unrecognised fifth by rendering it verbatim:
"verified" the probe got a 2xx — the key and the url agree "401" / "403" / any 3-digit decimal the upstream answered and REFUSED. The key, the url, or the pairing of the two is wrong. #3581's subject was a FastRouter key on an openrouter.ai url, which fails here identically to a mistyped key — that is the point. "unreachable" no response at all: dns, tls, egress refusal, timeout. A PROVIDER OUTAGE, not necessarily a bad credential, and never a reason to block a registration. "stale" a recorded verification was INVALIDATED because the endpoint's base_url or dialect changed after it was taken. It says "this was verified, but not against the endpoint you are looking at".
Empty means no probe outcome is recorded; pair it with last_verified_at. |
DeregisterLLMEndpointRequest
DeregisterLLMEndpointRequest soft-deletes an endpoint.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id names the endpoint. |
DeregisterLLMEndpointResponse
DeregisterLLMEndpointResponse confirms the soft-delete.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id echoes the deregistered endpoint. | |
| deregistered_at | google.protobuf.Timestamp | deregistered_at is the server time the row was soft-deleted. |
LLMEndpoint
LLMEndpoint is one registered upstream.
| Field | Type | Label | Description |
|---|---|---|---|
| id | string | id is the endpoint UUID. | |
| name | string | name is the operator-chosen label, unique among the org's ACTIVE endpoints (an endpoint that is disabled or deregistered releases its name). | |
| base_url | string | base_url is the egress destination. It has passed the full SSRF fence (LLD §8.3) at write time; see the validator for the refusal rules. | |
| dialect | string | dialect is a STABLE STRING: "anthropic" | |
| model_ids | string | repeated | model_ids holds model ids and/or id prefixes this endpoint serves. |
| min_clearance | int32 | min_clearance is the clearance floor (0–5) an agent must meet to use this endpoint. 0 means no floor. | |
| status | string | status is a STABLE STRING: "pending" | |
| version | int32 | version is bumped by the database whenever a materially significant field changes, and is what invalidates the gateway's resolution cache. Usable for optimistic-concurrency display. | |
| allow_unmetered | bool | allow_unmetered is the L5 acknowledgement that token metering on this endpoint is not trustworthy and budget enforcement is waived. Written ONLY by SetLLMEndpointUnmeteredAck. | |
| created_by | string | created_by is the member id of the operator who registered the endpoint. | |
| created_at | google.protobuf.Timestamp | created_at is when the endpoint was registered. | |
| updated_at | google.protobuf.Timestamp | updated_at is when the endpoint was last modified. | |
| auth_mode | string | auth_mode is a STABLE STRING: "key" |
"key" — the gateway resolves the tenant credential and presents it per the dialect's scheme. The default, and what every endpoint registered before this field existed carries. "none" — the gateway presents NO credential header. For a self-hosted keyless upstream (ollama, LiteLLM, vLLM).
THIS IS A STATEMENT ABOUT THE WIRE, NOT ABOUT TRUST. It does not widen the SSRF fence and it does not soften the dial-time egress check. Render an unknown future value verbatim rather than failing; the GATEWAY refuses one.
FOR RENDERERS (client#827, task 5C5): read this field BEFORE deciding what credential means. With auth_mode "key" an absent credential is broken; with "none" it is correctly configured. Those are two states, and a card that reads credential alone shows a healthy keyless endpoint as an error. |
| credential | CredentialSummary | | credential summarises the secret bound to this endpoint at the grain named by the list request's unit_id.
UNSET means NOT ASSESSED — the deployment has no vault configured, so presence is unknown. That is deliberately distinct from a SET message with present=false, which means "assessed, and there is no credential". |
| discovered_models | string | repeated | discovered_models is what the PROVIDER reported at GET {base_url}/{version}/models, captured by #3581's write-time credential probe in the same request (#3600).
── IT IS NOT WHAT THIS ENDPOINT ROUTES. READ model_ids FOR THAT. ───────
model_ids is the registered allowlist and the ONLY input to the gateway's model match. This field is the provider's catalogue, offered so the register/edit UI can present a PICKER instead of asking an operator to type model ids from the provider's docs — which is where the typos that produce a registered-but-routes-nothing endpoint come from.
A client MUST NOT infer access from it. An id here that is absent from model_ids is a model the tenant has NOT allowed, and rendering the two as one list would tell an operator they have access they do not have. Letting a provider's catalogue widen a tenant's allowlist would widen egress with nobody approving it.
EMPTY IS AMBIGUOUS ON THE WIRE AND discovered_at DISAMBIGUATES IT. proto3 cannot distinguish an absent repeated field from an empty one, so: discovered_at UNSET → never discovered. Render the typed list, no picker. discovered_at SET, list empty → the provider answered and serves NOTHING. That is a real finding; say so. |
| discovered_at | google.protobuf.Timestamp | | discovered_at is when discovered_models was captured, and it is the field that makes an empty catalogue readable (see above).
UNSET means NEVER DISCOVERED, which is not a failure: the endpoint predates the feature, the provider has no models route, it answered non-2xx, or no credential has been written yet. In all four the operator's typed model_ids stands unchanged.
It is refreshed wherever the verification probe runs — a credential set/rotate, and a base_url or dialect change — so it is also the answer to "how stale is this picker?". |
ListLLMEndpointsRequest
ListLLMEndpointsRequest selects one page of the caller org's endpoints.
| Field | Type | Label | Description |
|---|---|---|---|
| page_size | int32 | page_size caps the page. Defaults to 50 when unset or non-positive, and is clamped to 200. | |
| page_token | string | page_token is the opaque cursor returned as next_page_token by the previous call. Empty starts at the newest endpoint. A malformed token is InvalidArgument, never a silent restart from the beginning. | |
| unit_id | string | unit_id selects the org-unit grain for the embedded credential summary. Empty selects the org _default grain. A NON-EMPTY value returns Unimplemented in slice 1 — deliberately not a silent downgrade to _default, which would report a credential from a grain the caller did not ask about. The team grain lands in slice 2. |
ListLLMEndpointsResponse
ListLLMEndpointsResponse carries one page plus the cursor for the next.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoints | LLMEndpoint | repeated | endpoints is the page, newest first. |
| next_page_token | string | next_page_token is empty when no further pages remain. |
RegisterLLMEndpointRequest
RegisterLLMEndpointRequest creates an endpoint. There is deliberately NO credential field: the secret is written separately through AIGatewayEndpointCredentialService.
| Field | Type | Label | Description |
|---|---|---|---|
| name | string | name is the operator-chosen label. Must be non-empty and unique among the org's active endpoints; a duplicate is AlreadyExists. | |
| base_url | string | base_url is the egress destination, validated against the full SSRF fence (LLD §8.3). A refusal is InvalidArgument and names the rule that failed. | |
| dialect | string | dialect must be "anthropic" or "openai_compatible" in v1. | |
| model_ids | string | repeated | model_ids holds model ids and/or id prefixes. May be empty. |
| min_clearance | int32 | min_clearance is the clearance floor, 0–5. Out of range is InvalidArgument. | |
| auth_mode | string | auth_mode is "key" |
There is no *_set companion here and that is not an oversight: unlike min_clearance and model_ids, the zero value of this field is not a meaningful setting an operator might want to express. It is the absence of one. |
RegisterLLMEndpointResponse
RegisterLLMEndpointResponse returns the created endpoint, always in
status pending.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint | LLMEndpoint | endpoint is the newly created row. | |
| notices | string | repeated | notices are operator-facing messages about what the server did to the submitted values. THEY ARE NOT ERRORS: the registration succeeded and endpoint is the row that exists. |
A notice is emitted when the server STORED SOMETHING OTHER THAN WHAT WAS SENT, and today there is exactly one such case — the trailing /v1 strip (#3567). A silent normalisation is worse than no normalisation: the operator's own value is no longer what the registry holds, and nothing would tell them so.
A client MUST render every entry. An unrecognised notice is still a sentence an operator can read; there is deliberately no code to switch on, because a code invites a client to drop the ones it does not know. |
SetLLMEndpointStatusRequest
SetLLMEndpointStatusRequest drives activation and the kill-switch.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id names the endpoint. | |
| status | string | status is the target STABLE STRING: "pending" | |
| reason | string | reason is an optional operator note recorded in the audit detail. Never required; empty is accepted. |
SetLLMEndpointStatusResponse
SetLLMEndpointStatusResponse returns the endpoint after the transition.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint | LLMEndpoint | endpoint is the row at its new status. | |
| changed | bool | changed is false when the request named the status the endpoint was already in — the call was accepted and audited, but nothing moved. | |
| affected_capabilities | CapabilityImpact | repeated | affected_capabilities enumerates every context capability binding that references this endpoint — the blast radius of the kill-switch (MG-CTX.9 constraint 2, 5B2 #3155). |
WHY THIS NAMES RATHER THAN REFUSES. DeregisterLLMEndpoint REFUSES while a binding exists (5B1), because a soft-delete is not the operator's intent when five capabilities depend on the row. Disabling is different: it is MG-1.11's KILL-SWITCH, it is reversible, and it is the control an operator reaches for precisely when an endpoint is misbehaving. A kill-switch that can be refused is not a kill-switch. So the transition proceeds and the response — and the audit record — NAME what it took down.
Populated on the disabled transition, including the accepted no-op (an operator re-pressing the switch is entitled to the same answer). Empty on every other transition, and empty when nothing is bound.
READ IT ONLY TOGETHER WITH affected_capabilities_assessed. | | affected_capabilities_assessed | bool | | affected_capabilities_assessed distinguishes "nothing depends on this endpoint" from "the blast radius could not be determined".
This is CredentialSummary's NIL-means-NOT-ASSESSED rule, applied to a repeated field that has no nil. Collapsing the two would report a clean blast radius for a deployment that simply could not look — the exact substitution DeregisterLLMEndpoint's guard refuses outright.
It refuses outright and this one does NOT, and the asymmetry is deliberate: deregistration is destructive and may be refused, whereas this is MG-1.11's KILL-SWITCH and a kill-switch that can be refused is not a kill-switch. So the transition proceeds and the uncertainty is reported honestly instead.
False on every transition other than the disable, where the question was not asked at all. |
SetLLMEndpointUnmeteredAckRequest
SetLLMEndpointUnmeteredAckRequest waives budget enforcement for an endpoint whose token accounting cannot be trusted.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id names the endpoint. | |
| allow_unmetered | bool | allow_unmetered is the new value. False re-arms enforcement. | |
| justification | string | justification is REQUIRED and must be non-empty, on both the set and the clear. Turning off a spend control should cost the operator a sentence, and "who waived enforcement here, and why" must be answerable from a single audit predicate. |
SetLLMEndpointUnmeteredAckResponse
SetLLMEndpointUnmeteredAckResponse returns the endpoint after the write.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint | LLMEndpoint | endpoint is the row carrying the new allow_unmetered value. |
UpdateLLMEndpointRequest
UpdateLLMEndpointRequest is optional-by-presence: a field is applied only
when its *_set companion is true, because proto3 cannot distinguish an
unset scalar from its zero value (an important case here — min_clearance 0
and an empty model_ids list are both meaningful).
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id names the endpoint to edit. | |
| name | string | name is the replacement label. Applied only when name_set is true. | |
| name_set | bool | name_set marks name as present. | |
| base_url | string | base_url is the replacement egress destination, re-validated against the full SSRF fence. Applied only when base_url_set is true. | |
| base_url_set | bool | base_url_set marks base_url as present. | |
| dialect | string | dialect is the replacement dialect. Applied only when dialect_set is true. | |
| dialect_set | bool | dialect_set marks dialect as present. | |
| model_ids | string | repeated | model_ids is the replacement list — a full replace, not a merge. Applied only when model_ids_set is true, which is how an empty list is expressed. |
| model_ids_set | bool | model_ids_set marks model_ids as present. | |
| min_clearance | int32 | min_clearance is the replacement floor, 0–5. Applied only when min_clearance_set is true, which is how a floor of 0 is expressed. | |
| min_clearance_set | bool | min_clearance_set marks min_clearance as present. | |
| allow_unmetered | bool | allow_unmetered EXISTS ON THIS MESSAGE ONLY SO IT CAN BE REFUSED. Setting allow_unmetered_set on this L4 RPC returns InvalidArgument naming SetLLMEndpointUnmeteredAck, and the stored value is left unchanged. It is present rather than absent so that the bypass attempt is expressible, and therefore testable, rather than merely undocumented. | |
| allow_unmetered_set | bool | allow_unmetered_set marks allow_unmetered as present, and is the flag whose presence triggers the refusal above. | |
| auth_mode | string | auth_mode is the replacement mode, "key" |
Unlike RegisterLLMEndpointRequest, an EMPTY value here is InvalidArgument rather than a default to "key": setting auth_mode_set says the caller meant something, and defaulting it would silently rewrite a keyless endpoint into a keyed one on a request that asked for no such thing.
This IS editable at L4, unlike allow_unmetered. Switching between keyed and keyless is ordinary configuration, not a governance waiver — it waives no control, widens no egress, and each direction is caught by something else: "key" with no credential is credential_missing at request time, and every edit is audited naming the field. |
| auth_mode_set | bool | | auth_mode_set marks auth_mode as present. |
UpdateLLMEndpointResponse
UpdateLLMEndpointResponse returns the endpoint after the edit.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint | LLMEndpoint | endpoint is the updated row, with a bumped version when a cache-significant field changed. | |
| notices | string | repeated | notices are operator-facing messages about what the server did to the submitted values — same contract as RegisterLLMEndpointResponse.notices. |
Two cases reach here: the trailing /v1 strip (#3567), and the credential re-verification a base_url or dialect change triggers (#3581). The second one is why this field matters on the UPDATE path in particular: the recorded verification was taken against the OLD url, so a change invalidates it, and an operator who is not told that will read a stale "verified" badge as a statement about the endpoint they now have.
ON THE UPDATE PATH THE RE-VERIFICATION NOTICE PROMISES NOTHING ABOUT AN OUTCOME (#3596). The new base_url is UNCOMMITTED while this handler runs and the Model Gateway must read the row, so the probes cannot happen here: this response says how many verifications were marked stale and that re-verification is scheduled, and the badge updates when the post-commit sweep's RPCs land. A notice that reported an outcome would be reporting one taken against the PREVIOUS url — the exact confusion stale exists to end.
The list is deliberately open — a client renders every entry and switches on none — so a third producer needs no wire change. |
AIGatewayEndpointRegistryService
AIGatewayEndpointRegistryService is the tenant-facing registry of LLM upstreams the Model Gateway may dial on the tenant's behalf.
Every RPC is scoped to the caller's org, taken from the auth context and NEVER from the request body. Reads and writes run inside the request's RLS-bound scope transaction, so migration 203's tenant_isolation_llm_endpoints policy is the isolation mechanism; a cross-tenant id surfaces as NotFound.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| ListLLMEndpoints | ListLLMEndpointsRequest | ListLLMEndpointsResponse | ListLLMEndpoints returns one page of the caller org's endpoints, newest first, excluding deregistered (soft-deleted) rows. Requires L3. |
| RegisterLLMEndpoint | RegisterLLMEndpointRequest | RegisterLLMEndpointResponse | RegisterLLMEndpoint creates an endpoint in status pending. It NEVER accepts a secret, and it cannot create an already-approved endpoint — activation is a separate L5 decision. Requires L4. |
| UpdateLLMEndpoint | UpdateLLMEndpointRequest | UpdateLLMEndpointResponse | UpdateLLMEndpoint applies a partial edit. Fields are optional-by-presence via *_set companion flags. Rejects allow_unmetered with InvalidArgument — that field is L5-only, see SetLLMEndpointUnmeteredAck. Requires L4. |
| SetLLMEndpointStatus | SetLLMEndpointStatusRequest | SetLLMEndpointStatusResponse | SetLLMEndpointStatus moves an endpoint between pending, approved and disabled. Disabling is the org-wide kill-switch (MG-1.11) and takes effect on the next call, with no restart. Requires L5. |
| SetLLMEndpointUnmeteredAck | SetLLMEndpointUnmeteredAckRequest | SetLLMEndpointUnmeteredAckResponse | SetLLMEndpointUnmeteredAck records an operator acknowledgement that spend on this endpoint is not enforceable, and is the ONLY writer of allow_unmetered. justification is required and non-empty: this is the one write in slice 1 that turns off a spend control. Requires L5. |
| DeregisterLLMEndpoint | DeregisterLLMEndpointRequest | DeregisterLLMEndpointResponse | DeregisterLLMEndpoint soft-deletes an endpoint. It REFUSES with FailedPrecondition while any credential row exists for the endpoint — cleanup is explicit, never a cascade that silently drops secrets. Requires L4. |
upsquad/aigateway/v1/endpoint_binding.proto
ApproveLLMBindingRequest
ApproveLLMBindingRequest records the Manager's approval.
| Field | Type | Label | Description |
|---|---|---|---|
| binding_id | string | binding_id names the pending binding. | |
| reason | string | reason is an optional Manager note recorded on the row and in the audit detail. Never required on an approval; empty is accepted. |
ApproveLLMBindingResponse
ApproveLLMBindingResponse returns the approved binding.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | LLMBinding | binding is the row at approved. From this moment the resolver's join matches it and the binding is live. |
CreateLLMBindingRequest
CreateLLMBindingRequest attaches an endpoint to a unit.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id names the endpoint to bind. Must resolve under the caller's org (#2828). | |
| unit_id | string | unit_id names the org unit to bind it to. Must resolve under the caller's org (#2828). | |
| credential_mode | string | credential_mode is "team_shared" (default when empty) or "per_agent". | |
| require_per_agent | bool | require_per_agent forces the agent grain with no fallback. | |
| allow_translation | bool | allow_translation grants the dialect-translation capability (HLD §5.3). | |
| min_clearance | int32 | min_clearance is the binding's clearance floor, 0–5. Out of range is InvalidArgument. | |
| expected_rejected_binding_id | string | optional | Explicitly request a fresh pending binding after this exact binding was rejected. The server atomically retires only the matching rejected grant, preserving its decision history, then creates a new pending resource. Never retries approved/pending/withheld bindings and never approves the new request. Absent retains Create's existing duplicate-binding refusal. |
CreateLLMBindingResponse
CreateLLMBindingResponse returns the created binding, always pending.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | LLMBinding | binding is the newly created row. |
LLMBinding
LLMBinding is one llm_endpoint_units row, joined to its endpoint's name.
| Field | Type | Label | Description |
|---|---|---|---|
| id | string | id is the binding UUID. | |
| endpoint_id | string | endpoint_id is the bound llm_endpoints row. | |
| endpoint_name | string | endpoint_name is the endpoint's label, joined for display. A binding whose endpoint was deregistered still lists, and this field is what makes the row legible. | |
| unit_id | string | unit_id is the org unit this binding grants to. | |
| approval_status | string | approval_status is a STABLE STRING: "pending" |
NOTE THE SPELLING. llm_endpoint_units uses rejected where mcp_server_units uses denied (migration 212 §1 records the divergence deliberately). The two tables are siblings, not the same table. |
| binding_kind | string | | binding_kind is a STABLE STRING: "grant" | "withhold".
Withhold semantics are slice-3+, and the value is not writable through this service. It is reported because ANY READER MUST TREAT withhold AS DENY, never as unknown-therefore-allow — a client that renders it as a grant is reporting the opposite of the truth. |
| credential_mode | string | | credential_mode is a STABLE STRING: "team_shared" | "per_agent". |
| require_per_agent | bool | | require_per_agent forces the agent credential grain with no fallback to the team or org slot (MG-2.2). |
| allow_translation | bool | | allow_translation is the per-binding capability that permits a caller to opt into dialect translation on this endpoint (HLD §5.3, OQ-7). Default false at both levels; the team grants the capability, the caller exercises it. |
| min_clearance | int32 | | min_clearance is the binding's clearance floor (0–5). Gate 7a compares an agent's clearance against max(endpoint, binding).min_clearance, so this can only narrow access, never widen it below the endpoint's own floor. |
| requested_by | string | | requested_by is the member who created the binding — the actor whose request the Manager approves or rejects. Empty when unknown. |
| decided_by | string | | decided_by is the Manager who approved or rejected. Empty while pending. |
| decided_at | google.protobuf.Timestamp | | decided_at is when the decision was recorded. Unset while pending. |
| decision_reason | string | | decision_reason is the Manager's note on a rejection. May be empty. |
| configured_by | string | | configured_by is the member who attached the endpoint to the unit. |
| created_at | google.protobuf.Timestamp | | created_at is when the binding was requested. |
| updated_at | google.protobuf.Timestamp | | updated_at is when the binding last changed. |
| enabled_models | string | repeated | enabled_models is the team's NARROWING of what this endpoint serves (llm_endpoint_units.enabled_models, migration 262; core#3601).
EMPTY MEANS EVERY MODEL THE ENDPOINT SERVES IS ENABLED — the widest state, not the narrowest. That is the same emptiness rule the per-agent llm_enablement selection uses and the OPPOSITE of llm_endpoint_units itself, which is an enumerated grant whose absence grants nothing. Two columns of one row, two rules, each from that column's own semantics.
A non-empty list is exactly the enabled set, and every entry is one of the endpoint's own claim strings — so an entry may be a PREFIX (gpt-4) when the endpoint claims one, and it then enables every model under it, exactly as routing would.
A client that renders a checkbox grid MUST treat empty as all-ticked. Reading it as none-ticked would show a working team as having no model access. |
ListLLMBindingsRequest
ListLLMBindingsRequest selects one page of the caller org's bindings.
| Field | Type | Label | Description |
|---|---|---|---|
| unit_id | string | unit_id narrows the page to one org unit — the Manager's inbox view. Empty returns the whole org's bindings. A malformed uuid is InvalidArgument; a unit belonging to another org simply matches nothing, because the query is org-scoped. | |
| endpoint_id | string | endpoint_id narrows the page to one endpoint — "who can use this?". Empty does not narrow. Combines with unit_id as a conjunction. | |
| page_size | int32 | page_size caps the page. Defaults to 50 when unset or non-positive, and is clamped to 200. | |
| page_token | string | page_token is the opaque cursor returned as next_page_token by the previous call. Empty starts at the newest binding. A malformed token is InvalidArgument, never a silent restart from the beginning. |
ListLLMBindingsResponse
ListLLMBindingsResponse carries one page plus the cursor for the next.
| Field | Type | Label | Description |
|---|---|---|---|
| bindings | LLMBinding | repeated | bindings is the page, newest first, carrying EVERY approval status. |
| next_page_token | string | next_page_token is empty when no further pages remain. |
RejectLLMBindingRequest
RejectLLMBindingRequest records the Manager's rejection.
| Field | Type | Label | Description |
|---|---|---|---|
| binding_id | string | binding_id names the pending binding. | |
| reason | string | reason is the Manager's note. Optional, but it is the only thing that tells the requester WHY, so a client should prompt for it. |
RejectLLMBindingResponse
RejectLLMBindingResponse returns the rejected binding.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | LLMBinding | binding is the row at rejected. It stays visible on ListLLMBindings — a rejection the requester cannot see is a decision that never reached them. |
SetLLMBindingEnabledModelsRequest
SetLLMBindingEnabledModelsRequest replaces one binding's enabled-model set.
| Field | Type | Label | Description |
|---|---|---|---|
| binding_id | string | binding_id is the llm_endpoint_units row to narrow. It must be an approved GRANT binding under the caller's org. | |
| model_ids | string | repeated | model_ids is the FULL replacement set — this is a PUT, not a PATCH, so a client sends the state it wants and never a delta. A delta surface would make two concurrent saves silently compose into a third state nobody chose. |
EMPTY CLEARS THE NARROWING (every served model enabled). Every entry must be one of the endpoint's served claims. Duplicates are collapsed and the stored order is canonical, so two equivalent selections compare equal. |
SetLLMBindingEnabledModelsResponse
SetLLMBindingEnabledModelsResponse returns the binding as stored, so the client reads back the CANONICAL set rather than echoing its own request.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | LLMBinding | binding is the binding AS STORED, including the canonical enabled_models. The client reads this back rather than echoing its own request, so a set that equalled the endpoint's full served list is visibly reported as the empty (widest) state instead of two clients disagreeing about which one they saved. |
UpdateLLMBindingRequest
UpdateLLMBindingRequest is optional-by-presence: a field is applied only when
its *_set companion is true, because proto3 cannot distinguish an unset
scalar from its zero value — which matters here, since false and
min_clearance = 0 are both meaningful values.
| Field | Type | Label | Description |
|---|---|---|---|
| binding_id | string | binding_id names the binding to edit. | |
| credential_mode | string | credential_mode is the replacement grain policy. | |
| credential_mode_set | bool | credential_mode_set marks credential_mode as present. | |
| require_per_agent | bool | require_per_agent is the replacement value. | |
| require_per_agent_set | bool | require_per_agent_set marks require_per_agent as present, and is how false is expressed. | |
| allow_translation | bool | allow_translation is the replacement value. | |
| allow_translation_set | bool | allow_translation_set marks allow_translation as present, and is how false is expressed. | |
| min_clearance | int32 | min_clearance is the replacement floor, 0–5. | |
| min_clearance_set | bool | min_clearance_set marks min_clearance as present, and is how a floor of 0 is expressed. |
UpdateLLMBindingResponse
UpdateLLMBindingResponse returns the binding after the edit.
| Field | Type | Label | Description |
|---|---|---|---|
| binding | LLMBinding | binding is the updated row. Its approval_status is unchanged. |
AIGatewayEndpointBindingService
AIGatewayEndpointBindingService is the tenant-facing management surface for
llm_endpoint_units — which org units may use which registered LLM endpoint,
and on what terms.
Every RPC is scoped to the caller's org, taken from the auth context and NEVER from the request body. Reads and writes run inside the request's RLS-bound scope transaction, so migration 212's tenant_isolation_llm_endpoint_units policy is the isolation mechanism; a cross-tenant id surfaces as NotFound.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| ListLLMBindings | ListLLMBindingsRequest | ListLLMBindingsResponse | ListLLMBindings returns one page of the caller org's bindings, newest first, excluding detached (soft-deleted) rows. |
IT RETURNS EVERY APPROVAL STATUS — pending, approved AND rejected. That is the point of the RPC: it is the Manager's approval inbox as well as the team's roster. Requires L3. |
| CreateLLMBinding | CreateLLMBindingRequest | CreateLLMBindingResponse | CreateLLMBinding attaches an approved endpoint to an org unit, in status pending. It cannot create an already-approved binding — approval is a separate decision by the unit's Manager. Requires L4.
Both endpoint_id and unit_id must resolve under the caller's org (#2828); one that does not is NotFound, naming which of the two missed. |
| UpdateLLMBinding | UpdateLLMBindingRequest | UpdateLLMBindingResponse | UpdateLLMBinding edits the binding's write-surface knobs. It does NOT touch approval_status — editing the knobs is orthogonal to the approve/reject lifecycle. Requires L4 AND the caller must be the unit's Manager. |
| ApproveLLMBinding | ApproveLLMBindingRequest | ApproveLLMBindingResponse | ApproveLLMBinding transitions a pending binding to approved, which is the moment it stops being inert (MG-2.1). Terminal: a re-approve or an approve-after-reject is FailedPrecondition, never a silent no-op. Requires L4 AND the caller must be the unit's Manager. |
| RejectLLMBinding | RejectLLMBindingRequest | RejectLLMBindingResponse | RejectLLMBinding transitions a pending binding to rejected with an optional reason. Same terminal semantics and same authority as approve.
A rejected binding is NOT deleted: it stays visible on ListLLMBindings so the requester can see the decision and its reason. Requires L4 AND the caller must be the unit's Manager. | | SetLLMBindingEnabledModels | SetLLMBindingEnabledModelsRequest | SetLLMBindingEnabledModelsResponse | SetLLMBindingEnabledModels replaces the team's ENABLED-MODEL SET for one binding — the narrow-only model axis gate 7c enforces (core#3601's ruling §1/§3; migration 262).
NARROW-ONLY, AND NO SECOND APPROVAL
Binding approval (L4 AND the unit's Manager) approved THE CEILING: every model the endpoint serves. This write can only REMOVE from that ceiling, or restore up to it, so it needs no second approver — the same rule MG-2.6 gives SetAgentTools / SetAgentLLMEnablement one grain down. There is deliberately NO pending/approve axis on this set; the retired team_model_allowlists had one, and it is what let a team hold a PENDING entry for a model no bound endpoint served (core#3146).
AUTHORITY: L4+, OR L3+ WHEN THE CALLER IS THE BINDING UNIT'S MANAGER
Note this is NOT UpdateLLMBinding's flat L4 — it is the reduced bar the per-agent allowlist surface uses, because the unit's Manager is the authority over what her team may run. A unit with NO Manager is FailedPrecondition: no authority exists, so nobody at any clearance may narrow. The actor is taken from the VERIFIED SCOPE, never from a request field.
EMPTY MEANS EVERY SERVED MODEL, AND THE FULL SET IS STORED AS EMPTY
model_ids empty (the widest state) clears the narrowing: every model the endpoint serves is enabled. A list equal to the endpoint's full served set is CANONICALISED to the same stored NULL, so the UI's all-ticked state and its never-touched state are the same stored fact and two identical intents cannot produce two different rows.
WHAT IS REFUSED
- a binding that is not an approved GRANT: FailedPrecondition. Narrowing a pending, rejected or withhold binding is configuring something inert. - an entry the endpoint does not serve: InvalidArgument. The endpoint's own claims (and its discovered models) are the only legal entries, matched with the same matcher routing uses. - a narrowing that would leave NO model enabled: InvalidArgument. To remove every model, reject or withhold the binding — an endpoint bound but serving the team nothing is the dud state core#3146 measured. |
upsquad/aigateway/v1/endpoint_credential.proto
DeleteLLMEndpointCredentialRequest
DeleteLLMEndpointCredentialRequest identifies the credential slot to clear.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id is the llm_endpoints id whose credential is removed. Required. | |
| unit_id | string | unit_id scopes the delete to a team, identically to the Set request (UNIMPLEMENTED in slice 1). | |
| agent_id | string | agent_id scopes the delete to an agent, identically to the Set request (UNIMPLEMENTED in slice 1). Mutually exclusive with unit_id. |
DeleteLLMEndpointCredentialResponse
DeleteLLMEndpointCredentialResponse confirms the removal WITHOUT returning any secret material.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id echoes the endpoint the credential was removed for. | |
| vault_ref | string | vault_ref is the grain-qualified slot that was targeted. | |
| deleted | bool | deleted is true when a row was removed, false when none existed. False is a SUCCESS, not an error — the delete is idempotent. | |
| deleted_at | google.protobuf.Timestamp | deleted_at is the server time the delete was processed. |
LLMEndpointCredential
LLMEndpointCredential is one endpoint's credential slot at the requested grain.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id is the llm_endpoints id, parsed back out of the stored vault provider string "llm_endpoint_<endpoint_id>" for rows whose endpoint no longer exists. | |
| endpoint_name | string | endpoint_name is the endpoint's label, or EMPTY when the endpoint has been deregistered while its credential row survives. Such rows are still returned, deliberately, so the UI can offer cleanup rather than leaving an orphaned secret invisible. | |
| credential | CredentialSummary | credential is the masked summary for this slot. Its three states are normative (LLD §8.4): |
SET, present=true, hint → render the masked hint SET, present=false, "" → render "no credential set" UNSET → render "—", NEVER "no credential set"
The third state means NOT ASSESSED — no vault is configured, so presence could not be determined. Collapsing it into present=false would tell the operator a secret is missing when the server merely could not look. |
ListLLMEndpointCredentialsRequest
ListLLMEndpointCredentialsRequest selects the grain whose slots are listed.
| Field | Type | Label | Description |
|---|---|---|---|
| unit_id | string | unit_id lists the team grain (UNIMPLEMENTED in slice 1). Mutually exclusive with agent_id; both absent ⇒ the org _default grain. | |
| agent_id | string | agent_id lists the agent grain (UNIMPLEMENTED in slice 1). |
ListLLMEndpointCredentialsResponse
ListLLMEndpointCredentialsResponse returns one entry per endpoint, so the not-assessed state is expressible per slot.
| Field | Type | Label | Description |
|---|---|---|---|
| credentials | LLMEndpointCredential | repeated | credentials is one entry per live endpoint, plus one per orphaned credential row whose endpoint was deregistered. Oldest endpoint first. |
SetLLMEndpointCredentialRequest
SetLLMEndpointCredentialRequest carries the target endpoint and the raw secret. The tenant is taken from the auth context, NEVER from this body.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id is the llm_endpoints id the credential belongs to. Required, and MUST be an endpoint the caller's tenant owns — a cross-tenant or unknown id is NotFound. | |
| secret | string | secret is the raw upstream API key. Stored encrypted in the vault; NEVER echoed back, logged, or persisted to any other table. Write-only. | |
| unit_id | string | unit_id OPTIONALLY scopes the credential to one org_unit (team). Mutually exclusive with agent_id. Both absent ⇒ the org _default grain. |
SLICE 1: a non-empty value returns UNIMPLEMENTED naming llm_endpoint_units as the missing dependency. It is never silently downgraded to _default. |
| agent_id | string | | agent_id OPTIONALLY scopes the credential to one agent (team_id = "agent" + agent_id). Mutually exclusive with unit_id.
SLICE 1: also UNIMPLEMENTED. The agent grain needs no new table, but DeregisterLLMEndpoint's credential probe is ORG-GRAIN ONLY until slice 2 (LLD §8.1), so a writable agent grain now would let a deregistration drop an agent secret it never looked for — exactly the orphaning §8.1 warns about. The two widen together or not at all. |
SetLLMEndpointCredentialResponse
SetLLMEndpointCredentialResponse confirms the write WITHOUT returning the secret.
| Field | Type | Label | Description |
|---|---|---|---|
| endpoint_id | string | endpoint_id echoes the endpoint the credential was stored for. | |
| vault_ref | string | vault_ref is the resulting vault://provider_keys/<org>/<grain>/<provider> pointer. Opaque; it addresses a slot and carries no secret material. | |
| created | bool | created is true on a first-time set, false when this rotated an existing secret for the same (tenant, grain, endpoint). | |
| credential | CredentialSummary | credential is the post-write masked summary, read back from the vault's in-Postgres masked projection. Present=true and a hint on success. |
UNSET means NOT ASSESSED (LLD §8.4's third state): the write succeeded but the masked read-back could not be performed. It does NOT mean the write failed, and the FE must not render it as "no credential set". |
| updated_at | google.protobuf.Timestamp | | updated_at is the server time the credential was written. |
| notices | string | repeated | notices are operator-facing messages about the write (#3581). NOT ERRORS — the credential is stored and credential is the slot that now exists.
Two kinds arrive here, and the difference matters:
the PROBE RESULT, in words. credential.last_verify_status carries the machine-readable form; this is the sentence that says which half to go look at ("the key, or the url"), because a bare "401" sends an operator to re-paste a key that was never the problem (#3581's correction).
a PROVIDER-SHAPE HINT. Advisory only, never validation: the platform does not know every provider's key grammar and must not refuse on one. When the host is a provider whose documented key prefix is known and the pasted key does not match it, that is said here — and the key is stored regardless.
A failed probe does NOT fail this RPC and does not block approval. A provider outage must not stop a tenant from registering. |
AIGatewayEndpointCredentialService
AIGatewayEndpointCredentialService is the tenant's LLM-endpoint credential surface: set, rotate, remove, and a masked list.
Write-only for secret MATERIAL. Set is the only RPC that accepts a raw secret; no RPC returns the plaintext. Delete removes, never reveals. List returns masked summaries only.
CredentialSummary is NOT redeclared here — it is imported from
endpoint.proto, where LLD §8.4 placed it precisely so that the rule governing
a projection of a secret belongs to the projection rather than to whichever
service happens to return it. LLMEndpoint already embeds the same message.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| SetLLMEndpointCredential | SetLLMEndpointCredentialRequest | SetLLMEndpointCredentialResponse | SetLLMEndpointCredential stores the raw upstream API key for an endpoint in the tenant vault. Idempotent per (org, grain, endpoint): a second call ROTATES the secret and reports created=false. |
The response carries a CredentialSummary whose masked_hint is read back from the vault's IN-POSTGRES masked projection after the write — not computed in Go from the request. That keeps a single origin for every mask this platform renders, and means no code path in this service can produce a hint from plaintext held in process memory.
Clearance L4 (org and team grains). Emits exactly one llm_credential_set audit event carrying the endpoint, the grain and the actor — never the secret, and never the hint.
Errors: InvalidArgument (missing/malformed endpoint_id, empty or out-of-bounds secret, unit_id and agent_id both set), NotFound (unknown or cross-tenant endpoint), PermissionDenied (clearance), Unimplemented (a non-empty unit_id or agent_id in slice 1). | | DeleteLLMEndpointCredential | DeleteLLMEndpointCredentialRequest | DeleteLLMEndpointCredentialResponse | DeleteLLMEndpointCredential removes the stored credential at the given grain. IDEMPOTENT: deleting an absent credential SUCCEEDS with deleted=false — the FE's "remove" button must not error on a slot that was already cleared. Never returns secret material.
Clearance L4. Emits exactly one llm_credential_deleted audit event, including when deleted=false: an ATTEMPT on a credential slot is itself the auditable fact, and an unaudited attempt is indistinguishable from no attempt. |
| ListLLMEndpointCredentials | ListLLMEndpointCredentialsRequest | ListLLMEndpointCredentialsResponse | ListLLMEndpointCredentials enumerates the org's endpoints with a MASKED credential summary each, to render client#827's CredentialSlotCard.
SECURITY: masked hints only. The mask is derived inside Postgres (pgp_sym_decrypt is applied solely to extract the last 4 characters), so the full secret is never materialised in the process, a log, or the response.
ONE masked-list call per page, joined in Go — never N decrypting probes (LLD §8.4). The probe is a decrypting query; an N+1 on it is an N+1 on decryption.
Clearance L2 — a view-level read. Read-only: emits no audit event. |
upsquad/aigateway/v1/grain_budget.proto
DeleteGrainBudgetRequest
DeleteGrainBudgetRequest removes one grain's cap row.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain addresses the grain exactly as on GetGrainBudgetRequest. | |
| team_id | string | team_id is REQUIRED at BUDGET_GRAIN_TEAM and REFUSED at any other grain. | |
| agent_id | string | agent_id is REQUIRED at BUDGET_GRAIN_AGENT and REFUSED at any other grain. |
DeleteGrainBudgetResponse
DeleteGrainBudgetResponse carries the grain's state AFTER the removal.
| Field | Type | Label | Description |
|---|---|---|---|
| cap | GrainBudgetCap | cap is the grain's state after the delete — always BUDGET_CAP_STATE_UNCONFIGURED with budget_tokens absent. A caller that has just removed a control sees the control gone in the same response, without a second round trip that could observe a different instant. |
GetGrainBudgetRequest
GetGrainBudgetRequest addresses one grain.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain is REQUIRED. BUDGET_GRAIN_UNSPECIFIED is refused, and an undeclared number is refused — proto3 enums are OPEN and a defaulted grain answers a different question indistinguishably. |
BUDGET_GRAIN_ORG is REFUSED here, and the refusal names why: the org cap is feature_gates.limit_value at the org's PLAN TIER. It is a property of the subscription, not a tenant-settable number, and a Set that appeared to accept it would be a billing change wearing a budget control's clothes. |
| team_id | string | | team_id is REQUIRED at BUDGET_GRAIN_TEAM and REFUSED at any other grain. It is an org_units.id — the same identifier team_budgets.team_id holds and the same one llm_endpoint_units.unit_id joins on. |
| agent_id | string | | agent_id is REQUIRED at BUDGET_GRAIN_AGENT and REFUSED at any other grain. An agents.id belonging to the caller's org. |
GetGrainBudgetResponse
GetGrainBudgetResponse carries the grain's state.
The three RPCs return three DISTINCT message types wrapping one cap, rather
than sharing GrainBudgetCap directly. That is buf's one-response-type-per-RPC
rule, and it is worth having beyond lint compliance: it leaves room for a
future response to carry something only one of the three needs — a Set that
reports the previous value, say — without a breaking change or a shared
message accreting fields two thirds of its readers must ignore.
| Field | Type | Label | Description |
|---|---|---|---|
| cap | GrainBudgetCap | cap is the grain's state as read. | |
| can_write | bool | optional | can_write projects the SAME authority gate used by Set/Delete for an otherwise valid request. Absent means an older server has no projection; explicit false is a current denial. Writes always recheck authority, and true does not promise valid token values or successful storage operations. |
| denial_reason | string | denial_reason is the server's named authority denial verbatim when can_write is false; empty when true. A Manager lookup that finds no live unit in the caller's organisation is a denial: "grainbudget: no such org unit in this organisation". The readable cap is retained; Set/Delete still return NotFound. Other NotFound errors and storage failures fail the RPC instead of being represented as an authority decision. |
GrainBudgetCap
GrainBudgetCap is what all three RPCs return: the state of one grain AFTER the call.
Delete returns it too, reporting BUDGET_CAP_STATE_UNCONFIGURED, rather than returning nothing. A caller that has just removed a control should be able to see that the control is gone from the same response that removed it, without a second round trip that could observe a different instant.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain echoes the grain that was addressed, resolved by the SERVER. | |
| bucket_key | string | bucket_key names WHAT was answered: the team id at team grain, the agent id at agent grain. Never the org — the org is the caller and is not echoed to a request that could not name it. | |
| cap_state | BudgetCapState | cap_state is BUDGET_CAP_STATE_IN_FORCE when a cap row exists, BUDGET_CAP_STATE_UNCONFIGURED when it does not, and BUDGET_CAP_STATE_UNREADABLE when no billing window could be derived for the org (which is UNKNOWN, never unlimited — #2861). |
BUDGET_CAP_STATE_LAPSED IS NEVER RETURNED BY THIS SERVICE. The windows recur, so a present row is always in force. The member remains in the enum because proto3 cannot remove one and because the burndown's contract still documents it; a client must still handle it as "not capping right now". |
| budget_tokens | int64 | optional | budget_tokens is the cap now in force. ABSENT when cap_state is not IN_FORCE — an absent field cannot be mistaken for a cap of zero, which is the one confusion this surface refuses at the input as well. |
| window_start | google.protobuf.Timestamp | | window_start / window_end are the half-open billing period the cap is measured over — DERIVED from the org's billing anchor, identical to the window gate 7d sums usage over and to the one the burndown reports. One source: a second derivation is exactly what this change removed.
Both are ABSENT when no window could be derived (cap_state UNREADABLE). |
| window_end | google.protobuf.Timestamp | | window_end is the EXCLUSIVE end of that window. Same presence rules as window_start; both are absent together, never one alone, because a half-derived window bounds nothing. |
SetGrainBudgetRequest
SetGrainBudgetRequest sets one grain's token cap.
NOTE WHAT IS ABSENT: no period, no budget_usd, no alert_pct /
throttle_pct, no org. Each absence is argued in this file's header, and the
period absence is an acceptance criterion asserted from the descriptor.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | BudgetGrain | grain addresses the grain exactly as on GetGrainBudgetRequest, with the same refusals. BUDGET_GRAIN_ORG is refused: the org cap is a property of the subscription's plan tier, not a tenant-settable number. | |
| team_id | string | team_id is REQUIRED at BUDGET_GRAIN_TEAM and REFUSED at any other grain. | |
| agent_id | string | agent_id is REQUIRED at BUDGET_GRAIN_AGENT and REFUSED at any other grain. | |
| budget_tokens | int64 | budget_tokens is the cap for one billing period. MUST be > 0. |
It is a plain int64 rather than optional, and that is load-bearing on a SET: an absent optional and a zero are both "the operator did not give me a number", and this surface must be able to refuse that rather than treat it as "leave the cap alone" (which would make a typo a no-op) or as "no cap" (which would make a typo a removal). |
SetGrainBudgetResponse
SetGrainBudgetResponse carries the grain's state AFTER the write.
| Field | Type | Label | Description |
|---|---|---|---|
| cap | GrainBudgetCap | cap is the grain's state after the upsert, read back through the same projection Get uses. It is a re-read rather than an echo of the request: the window is derived and the request could not name it, so echoing would report a cap the server never confirmed. |
AIGatewayGrainBudgetService
AIGatewayGrainBudgetService is the tenant's surface for the two token caps gate 7d composes alongside the org's plan limit.
THE ORG IS NEVER NAMEABLE BY A REQUEST
No message in this file declares an org field. The org comes from the
authenticated caller, from the same mechanism that yields the clearance, and
TestGrainBudget_NoRequestMessageDeclaresAnOrgField derives that from the
descriptors. A tenant-supplied org on a WRITE surface is a cross-tenant
mutation, which is a strictly worse failure than the cross-tenant read the
same guard prevents on the burndown.
AUTHORISATION, AND WHY IT IS ASYMMETRIC BETWEEN THE GRAINS
GetGrainBudget L3
SetGrainBudget L4 + the unit's Manager at TEAM grain
DeleteGrainBudget L4 + the unit's Manager at TEAM grain
TEAM grain is Manager-gated because org_units.manager_member_id is the
authority over a unit — the same authority ADR-0021 gave binding approval, and
the pattern #2928 asked for.
AGENT grain is L4 WITHOUT a Manager gate, and the asymmetry is deliberate
rather than an oversight: agents carries no org-unit membership, so there is
no manager row to consult, and inventing one would put the agent cap out of
reach of the product. The blast radius is bounded by the composition itself —
min-of means an agent cap can only ever NARROW what the team and org caps
already allow, so raising one cannot widen a tenant's effective spend. The
agent is still constrained to the caller's org STRUCTURALLY, in the INSERT.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| GetGrainBudget | GetGrainBudgetRequest | GetGrainBudgetResponse | GetGrainBudget reports the cap in force at one grain and the window it is measured over. L3 — the same floor as the burndown, for the same reason: a team's own cap must not be out of reach of the person who runs that team. |
| SetGrainBudget | SetGrainBudgetRequest | SetGrainBudgetResponse | SetGrainBudget creates or replaces the token cap at one grain. |
It is an UPSERT on budget_tokens and on nothing else. On an existing row it touches exactly one column; period_start / period_end are not in the SET clause, which is the architect's AC discharged in SQL rather than in a handler.
budget_tokens MUST be > 0. Zero is REFUSED with InvalidArgument rather than accepted: budget.PGConfigSource and GrainBudget.inForce both read a non-positive cap as NO CAP, so an operator typing 0 meaning "stop this team" would get "this team is uncapped" — the fail-open direction, silently. DeleteGrainBudget is how a grain is returned to uncapped, and it says so. |
| DeleteGrainBudget | DeleteGrainBudgetRequest | DeleteGrainBudgetResponse | DeleteGrainBudget removes the cap row, returning the grain to UNCAPPED.
It exists so that capping is not a one-way door. Without it an operator who capped a team by mistake could only ever raise the number, never withdraw the control, and "uncapped" — the state every team is in today, and the state Compose documents as admitted — would be unreachable through the product for any team that had ever been capped.
Removing a spend control is audited under its OWN action, llm_grain_budget_cleared, not as a set with an absent value. |
upsquad/aigateway/v1/models_catalog.proto
AIGatewayCatalogConfig
AIGatewayCatalogConfig is the combined Get/Update payload.
| Field | Type | Label | Description |
|---|---|---|---|
| providers | Provider | repeated | Providers is the full platform provider list. |
| models | Model | repeated | Models is the full platform model list. |
| tenant_overrides | TenantOverride | repeated | TenantOverrides is the caller's tenant-scoped overrides on top of the catalog. Server-stamped; entries for other tenants are never returned. |
| updated_at | google.protobuf.Timestamp | UpdatedAt is stamped server-side on every Update. Never accepted from the client. | |
| version | string | Opaque concurrency token for the platform catalogue plus this tenant's overrides. Echo through Update.expected_version. Unlike updated_at this survives deleting the final override. Never parse or manufacture it. |
GetAIGatewayCatalogConfigRequest
GetAIGatewayCatalogConfigRequest is empty — tenant id is sourced from the authenticated context.
GetAIGatewayCatalogConfigResponse
GetAIGatewayCatalogConfigResponse carries the merged catalog.
| Field | Type | Label | Description |
|---|---|---|---|
| config | AIGatewayCatalogConfig | Config is the merged catalog: platform providers + platform models + the caller's tenant overrides. Tenant overrides for other tenants are never returned. |
Model
Model is a registered model with its token budget and provider binding. Platform-global — tenant overrides live in TenantOverride.
| Field | Type | Label | Description |
|---|---|---|---|
| id | string | ID is the on-the-wire model id (e.g. "claude-opus-4-7"). | |
| provider_id | string | ProviderID is the foreign key into Provider.id. Validation rejects unknown provider ids. | |
| display_name | string | DisplayName is the human-facing label. | |
| max_input_tokens | int64 | MaxInputTokens is the compile-time context-window ceiling. 0 means "no ceiling — defer to provider". | |
| max_output_tokens | int64 | MaxOutputTokens is the per-call output ceiling (tokens). 0 means "no ceiling". |
Provider
Provider is a platform-registered LLM provider. Global — all tenants see the same list. Enabled=false hides the provider's models from the runtime router without deleting the row.
| Field | Type | Label | Description |
|---|---|---|---|
| id | string | ID is the stable identifier, lowercase dotted (e.g. "anthropic", "openai"). Immutable primary key. | |
| display_name | string | DisplayName is the human-facing label rendered in the catalog UI. | |
| enabled | bool | Enabled gates every model owned by this provider. Flipping to false is effectively a platform-wide disable. | |
| default_base_url | string | DefaultBaseURL overrides the provider SDK's hard-coded base URL when non-empty. Used for private-region routes or self-hosted compatible gateways. |
TenantOverride
TenantOverride is a per-tenant tweak on top of the platform catalog. tenant_id is server-stamped from the authenticated context — never accepted from the wire.
| Field | Type | Label | Description |
|---|---|---|---|
| model_id | string | ModelID must match a Model.id in the catalog. A row referencing a missing model is rejected at Update time. | |
| enabled | bool | Enabled hides the model from the caller's tenant when false, regardless of the platform-level Model/Provider state. | |
| rate_limit_rpm | int32 | RateLimitRPM caps requests-per-minute for this tenant on this model. 0 means "use the platform default". |
UpdateAIGatewayCatalogConfigRequest
UpdateAIGatewayCatalogConfigRequest carries the desired new catalog. L5 actors see providers + models applied; L4 actors see only tenant_overrides applied (providers / models are ignored with a server-side log line, not an error, to keep the round-trip simple).
| Field | Type | Label | Description |
|---|---|---|---|
| expected_version | string | Required token returned by Get. Missing => FailedPrecondition; stale => Aborted without a write. Re-read and reconcile before retrying. This applies to L4 overrides and L5 platform writes; there is no unconditional-write arm. | |
| config | AIGatewayCatalogConfig | Config is the desired new catalog. L5 actors may mutate providers + models + tenant_overrides; L4 actors mutate tenant_overrides only (platform fields on the wire are dropped server-side). |
UpdateAIGatewayCatalogConfigResponse
UpdateAIGatewayCatalogConfigResponse echoes the merged catalog after the write.
| Field | Type | Label | Description |
|---|---|---|---|
| config | AIGatewayCatalogConfig | Config is the merged catalog after the write, same shape as Get. |
AIGatewayCatalogConfigService
AIGatewayCatalogConfigService exposes the providers / models catalog as a single Get/Update pair. LLD-4 §2.1 shape.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| GetAIGatewayCatalogConfig | GetAIGatewayCatalogConfigRequest | GetAIGatewayCatalogConfigResponse | Get the merged catalog (platform providers + platform models + caller's tenant overrides). Read-only, L3. |
| UpdateAIGatewayCatalogConfig | UpdateAIGatewayCatalogConfigRequest | UpdateAIGatewayCatalogConfigResponse | Update replaces the catalog for the caller's scope: - L5 actor: full catalog (providers + models + tenant overrides) - L4 actor: tenant overrides only; providers / models ignored Emits one config.updated audit event on success. |
upsquad/aigateway/v1/provider_key.proto
SetProviderKeyRequest
SetProviderKeyRequest carries the provider id + the raw secret. The tenant is taken from the auth context (never from the request body).
| Field | Type | Label | Description |
|---|---|---|---|
| provider | string | provider is the lowercase provider id, e.g. "anthropic" | |
| api_key | string | api_key is the raw provider secret. It is stored encrypted in the vault and is NEVER echoed back, logged, or persisted to any table. | |
| unit_id | string | unit_id selects the TEAM grain: the key is stored under team_id=<unit_id> and the AI Gateway prefers it over the org key for calls made by that team. Must be a live org_unit of unit_type 'team' IN THE CALLER'S TENANT; an unknown, deleted or foreign id is PERMISSION_DENIED, indistinguishably. Mutually exclusive with agent_id. | |
| agent_id | string | agent_id selects the AGENT grain (team_id="agent<agent_id>"). Must be a live agent IN THE CALLER'S TENANT, same denial rule as unit_id. Mutually exclusive with unit_id. |
SetProviderKeyResponse
SetProviderKeyResponse confirms the write WITHOUT returning the secret.
| Field | Type | Label | Description |
|---|---|---|---|
| provider | string | provider echoes the stored provider id. | |
| key_hint | string | key_hint is a masked fingerprint (last 4 chars only), safe to render in the UI to confirm which key was stored — e.g. "••••AB12". | |
| created | bool | created is true when this was a first-time set, false when it rotated an existing key for the same (tenant, provider). | |
| updated_at | google.protobuf.Timestamp | updated_at is the server time the key was written. | |
| grain | string | grain echoes WHICH SLOT was written: "org" |
AIGatewayProviderKeyService
AIGatewayProviderKeyService is the thin tenant BYOK key-write surface. Write-only: there is deliberately no Get/List RPC — secret material never leaves the vault after it is set.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| SetProviderKey | SetProviderKeyRequest | SetProviderKeyResponse | SetProviderKey upserts the caller tenant's API key for an LLM provider at one of three grains. Idempotent on (org, grain slot, provider): a second call for the same slot rotates that slot's key and leaves the other two untouched. Emits one provider_key.set audit event on success carrying only the provider, the grain and a masked hint — never the raw key. Requires L4 clearance at every grain. |
upsquad/aigateway/v1/usage.proto
ExportUsageRequest
ExportUsageRequest selects one page to render as CSV. Same selection as QueryUsageRequest; a separate message because the two RPCs' requests must stay independently evolvable.
| Field | Type | Label | Description |
|---|---|---|---|
| filter | UsageFilter | filter is the grain, window and narrowing predicates. Required. | |
| page_size | int32 | page_size bounds the page. Defaults to 50, capped at 200. | |
| page_token | string | page_token continues a previous page, exactly as on QueryUsage. The two RPCs mint and accept the same token format, so a client may page a query and export the same page. |
ExportUsageResponse
ExportUsageResponse carries the CSV document for one page.
| Field | Type | Label | Description |
|---|---|---|---|
| csv | bytes | csv is the rendered page: a header row followed by one row per bucket, RFC 4180, UTF-8, no BOM. |
The columns are derived from UsageRow's own descriptor and a CI guard fails when a field is added to UsageRow without a column (LLD #2901 §4.4), so the CSV cannot silently drop a dimension the structured API reports.
AN UNKNOWN COST IS AN EMPTY CELL, NEVER "0". Same rule as the absent cost_usd field, one encoding down: a spreadsheet sums an empty cell as nothing and a "0" as nothing, but only the empty cell survives the round trip as "unknown" to a human reading it. |
| filename | string | | filename is a suggested download name, e.g. "upsquad-usage-model-20260901-20260903.csv". Advisory only. |
| totals | UsageTotals | | totals covers the whole filtered window, structured rather than appended to the CSV body — see the RPC comment. |
| next_page_token | string | | next_page_token is empty when this is the last page. |
LatencyDistribution
LatencyDistribution is percentile_cont over one latency column, for one
population, TOGETHER WITH THE DENOMINATOR IT WAS COMPUTED OVER.
WHY THIS IS A MESSAGE AND NOT THREE FIELDS ON UsageAggregate
Somebody will propose folding these three numbers into UsageAggregate as a
simplification. It is not one, and the reason is arithmetic rather than
taste:
UsageAggregate IS THE SUMMED SHAPE. It is declared ONCE and used by both
UsageRow and UsageTotals precisely so that a per-bucket counter and its
total cannot drift apart — the server produces `UsageTotals.aggregate` by
RE-AGGREGATING the buckets. Every field on it is therefore a field that
gets ADDED UP.
A SUMMED PERCENTILE IS NOT A PERCENTILE. p99 of a union is not the sum of
the parts' p99s. It is not their average, their maximum, or any other
function of them: bucket-level percentiles do not carry the information
needed to reconstruct the window-level one. There is NO arithmetic that
recovers it, so there is no way to make the shared shape correct.
The failure would be SILENT AND CONFIDENT. A `p99_ms` on `UsageAggregate`
would render on the totals row as an ordinary number with no error, no
NULL, and no way for a client to tell it apart from a measured one.
So the window-level distribution below is not derived from the page at all.
QueryLatency runs a SECOND percentile_cont over the whole filtered
window's raw rows. Two independent computations over two populations, never
one composed from the other — which is exactly the thing a shared summed
message cannot express.
THE PERCENTILES ARE DECIMAL STRINGS, NEVER A double
Same rule as cost_usd, one column over. PostgreSQL's percentile_cont
returns double precision when ordering an integer column, and it shows: on
this repo's own probe, p90 over {10,20,30,1000} came back as
709.0000000000002. That float artefact is not a measurement — the source
column is whole milliseconds — and putting it on the wire as a JSON number
invites every consumer to re-inherit it.
The server therefore quantises to a DECLARED SCALE of 3 decimal places (microsecond resolution, three orders of magnitude finer than the source column's own quantisation) and sends the decimal text PostgreSQL produced. The string is the value; there is no more precision behind it.
ABSENT IS NOT ZERO, AND HERE THAT IS THE WHOLE POINT
When sample_count == 0 the three percentile fields are ABSENT. A bucket in
which nothing was measured has an UNKNOWN latency, not a fast one, and 0 is
the single most dangerous value a reliability panel could receive: it renders
green.
| Field | Type | Label | Description |
|---|---|---|---|
| p50_of_measured_ms | string | optional | p50_of_measured_ms is percentile_cont(0.50) OVER THE MEASURED EVENTS ONLY, in milliseconds, as a decimal string at scale 3. |
WHY THE NAME CARRIES of_measured, AND WHY IT IS NOT p50_ms
The task that specified this message called the field p50_ms. That name asserts something the value does not have whenever unmeasured_events > 0: a reader takes p50_ms to be the median of THIS BUCKET'S CALLS, and it is the median of the SUBSET that carried a timing.
The correction is the one #3029 landed hours earlier, where upper_bound_usd became modelled_ceiling_usd: THE FIELD NAME IS THE ONE THING EVERY CONSUMER READS. The alert rule, the CSV header, the dashboard binding and the SQL over an export all read the name; the comment, the sibling counter and the denominator all sit where a consumer may never look. An honest value under a name that overstates it is the same defect whether the subject is money or milliseconds.
IT IS NOT A BOUND IN EITHER DIRECTION
Stated because the obvious repair — "treat it as optimistic" — is also wrong. Nothing is known about the unmeasured events' latency: they may be faster or slower than the measured ones, so the full-population percentile may sit above OR below this figure. It supports exactly one conclusion: "among the sample_count calls that were timed, half were at or under this". The parent issue makes the same point about cache-hit ratio, which is a ratio of two lower bounds and therefore not itself a bound.
ABSENT when sample_count == 0 — see the message comment. A client that renders absent as 0 reports an instantaneous response that never happened. |
| p90_of_measured_ms | string | optional | p90_of_measured_ms is the 90th percentile over the measured events, same encoding, same absence rule and the same naming argument as p50. |
| p99_of_measured_ms | string | optional | p99_of_measured_ms is the 99th percentile over the measured events, same encoding, same absence rule and the same naming argument as p50.
This is the field the founder's constraint names directly: a p99 over the rows that happened to have timings is not a p99, and it is worse than no number because it looks like one. It is safe to publish only because the name says which population it covers and unmeasured_events says how much of the bucket it does not. |
| sample_count | int64 | | sample_count is how many events in this population carried a USABLE measurement and therefore contributed to the percentiles above.
NOTE WHAT A SMALL VALUE DOES TO p99. percentile_cont interpolates, so p99 over four samples is a statement about those four samples and about nothing else. This count travels beside the percentiles so the client can decide whether the number deserves to be rendered at all; the server imposes no minimum, because a threshold in the server is a place for a silently withheld answer to hide. |
| unmeasured_events | int64 | | unmeasured_events is how many events in the SAME population carried no usable measurement and were therefore excluded from the percentiles above.
THIS FIELD IS THE HONESTY AXIS AND IT IS NOT DECORATION
A percentile over a partially-measured population MUST state its denominator. A p99 computed over "the rows that happened to have a timing" is not the window's p99 — and it is WORSE than no number, because it looks like one. sample_count and unmeasured_events together are that denominator: they PARTITION the bucket's calls, and the server refuses to answer when they do not (LatencyRow.calls).
It is a SEPARATE COUNTER for the same reason unknown_cost_events is: no arithmetic on the percentiles can absorb it. There is no field it could be folded into.
THE TWO SPELLINGS OF "NOT MEASURED", BOTH COUNTED HERE
Migration 227's COMMENT ON COLUMN llm_usage_effective_cost.total_ms is normative and it names both:
NULL no observation was ever recorded — pre-018 rows and the two worker-side producers; 0 the WRITER'S OWN sentinel for "not measured", stored as a real integer (internal/aigateway/handler/handler.go:566, "Both return 0 for not measured").
A percentile that excluded only the NULLs would silently admit the second spelling as a measured 0 ms — dragging p50 toward zero AND inflating sample_count, so the denominator would lie in the reassuring direction. Both spellings are excluded from the statistic and counted here.
ttft_ms IN PARTICULAR IS EXPECTED TO BE MOSTLY UNMEASURED
Time-to-first-token is meaningless for a non-streaming call and 227 records that it is 0 there. So on a corpus of non-streaming traffic this field legitimately equals calls and sample_count is 0 — which reports "we do not know", correctly and loudly, instead of "0 ms". |
LatencyRow
LatencyRow is one bucket's latency shape at the requested grain.
bucket_key carries the same opaque, per-grain semantics as UsageRow's,
including EMPTY BEING A REAL BUCKET — see UsageRow.bucket_key.
| Field | Type | Label | Description |
|---|---|---|---|
| bucket_key | string | bucket_key is the grain's value. Opaque; empty is the unattributed bucket. | |
| calls | int64 | calls is the number of ledger events in the bucket — THE DENOMINATOR both distributions below partition. |
It is carried explicitly rather than left implicit in sample_count + unmeasured_events because the server checks that identity on every row before answering, and a check whose two sides come from one number is not a check. calls, each sample_count and each unmeasured_events are independent aggregates in SQL. |
| total_ms | LatencyDistribution | | total_ms is the wall-clock distribution over this bucket. |
| ttft_ms | LatencyDistribution | | ttft_ms is the time-to-first-token distribution over this bucket. Expect a large unmeasured_events here on non-streaming traffic — see that field. |
LatencyTotals
LatencyTotals is the distribution over the WHOLE filtered window, independent of pagination.
IT IS RECOMPUTED, NOT ROLLED UP
The percentiles here come from a SECOND percentile_cont over every raw row
the filter matched — not from the rows on the page, and not from any function
of their percentiles, because no such function exists. The counts (calls,
sample_count, unmeasured_events) DO sum, and the server checks that they
do.
Always present, including on an empty window, where every counter is zero and all six percentiles are ABSENT.
| Field | Type | Label | Description |
|---|---|---|---|
| buckets | int64 | buckets is how many distinct grain keys the filter matched across every page — the same dimension UsageTotals.buckets carries, for the same reason. | |
| calls | int64 | calls is the event count over the whole window: the denominator both distributions below partition. | |
| total_ms | LatencyDistribution | total_ms is the wall-clock distribution over the whole window. | |
| ttft_ms | LatencyDistribution | ttft_ms is the time-to-first-token distribution over the whole window. |
QueryLatencyRequest
QueryLatencyRequest selects one page of latency distributions. Same selection as QueryUsageRequest; a separate message so the two RPCs' requests stay independently evolvable.
| Field | Type | Label | Description |
|---|---|---|---|
| filter | UsageFilter | filter is the grain, window and narrowing predicates. Required, and validated identically to QueryUsage's — UNSPECIFIED and any unmapped grain number are REFUSED, never defaulted. | |
| page_size | int32 | page_size bounds the page. Defaults to 50, capped at 200. | |
| page_token | string | page_token continues a previous page, exactly as on QueryUsage. |
QueryLatencyResponse
QueryLatencyResponse is one page of buckets plus the window's distribution.
| Field | Type | Label | Description |
|---|---|---|---|
| rows | LatencyRow | repeated | rows is the page, ordered by bucket_key ascending — same ordering and same reasoning as QueryUsageResponse.rows. |
| totals | LatencyTotals | totals covers the whole filtered window, recomputed rather than rolled up. | |
| next_page_token | string | next_page_token is empty when this is the last page. |
QueryUsageRequest
QueryUsageRequest selects one page of aggregated usage.
| Field | Type | Label | Description |
|---|---|---|---|
| filter | UsageFilter | filter is the grain, window and narrowing predicates. Required. | |
| page_size | int32 | page_size bounds the page. Defaults to 50, capped at 200. | |
| page_token | string | page_token continues a previous page. Opaque; pass back QueryUsageResponse.next_page_token verbatim. A token minted for a DIFFERENT filter is refused rather than silently re-anchored. |
QueryUsageResponse
QueryUsageResponse is one page plus the window's totals.
| Field | Type | Label | Description |
|---|---|---|---|
| rows | UsageRow | repeated | rows is the page, ordered by bucket_key ascending. |
ASCENDING KEY, NOT DESCENDING COST. Keyset pagination needs a stable, unique ordering column; cost is neither unique nor stable while the ledger is being appended to, and a cost-ordered cursor silently skips and repeats buckets between pages. A dashboard that wants "top spenders" sorts a page, or asks for a window small enough to fit one. | | totals | UsageTotals | | totals covers the whole filtered window, not this page. | | next_page_token | string | | next_page_token is empty when this is the last page. |
UsageAggregate
UsageAggregate is the counted shape of one bucket. ONE declaration, used by both UsageRow and UsageTotals, so a per-row counter and its total cannot come to mean different things.
EVERY FIELD HERE SUMS. THAT IS WHY NO PERCENTILE MAY EVER BE ADDED TO IT.
The sharing above is the whole design: UsageTotals.aggregate is produced by
re-aggregating the buckets, so a field on this message is a field the server
will ADD UP across buckets. That is correct for a count and for a sum, and it
is WRONG for a percentile — p99 of a union is not the sum of the parts' p99s,
and it is not their average either. No arithmetic on bucket-level percentiles
reconstructs the window-level one, so a p99_ms here would make the totals
row fabricate a number that reads exactly like a measured one.
Latency percentiles therefore live on their own message, LatencyDistribution,
served by their own RPC that recomputes them from the raw rows at each level.
See that message's comment for the full argument, including why folding it
back into this one is not a simplification.
| Field | Type | Label | Description |
|---|---|---|---|
| calls | int64 | calls is the number of ledger events in the bucket. It always equals priced_events + unknown_cost_events; the server asserts that partition on every row and on the totals before answering. | |
| input_tokens | int64 | input_tokens is the summed prompt tokens. | |
| output_tokens | int64 | output_tokens is the summed completion tokens. | |
| cache_creation_tokens | int64 | cache_creation_tokens is the summed cache-write tokens (migration 203). | |
| cache_read_tokens | int64 | cache_read_tokens is the summed cache-read tokens (migration 203). | |
| cost_usd | string | optional | cost_usd is the summed cost of the PRICED events only, as a decimal string at the column's own NUMERIC(14,8) scale. |
ABSENT means "no priced event in this bucket" and MUST NOT be rendered as zero: a bucket of unpriced calls costs an unknown amount, not nothing. See the file header for why absent-vs-zero is the only input on which the honest and the fabricating query disagree. |
| priced_events | int64 | | priced_events is how many of calls contributed to cost_usd. |
| unknown_cost_events | int64 | | unknown_cost_events is how many of calls had NO cost (OQ-3's separate bucket). It is NEVER summed into cost_usd — there is no field it could be summed into, which is the point of carrying it as a count.
A non-zero value means cost_usd is a LOWER BOUND, and a client that renders it as the total is under-reporting by an unknown amount. |
| metered_events | int64 | | metered_events is how many of calls carry metering_confidence = 'complete'. ONLY these calls' token counts are totals; see unmetered_events. |
| unmetered_events | int64 | | unmetered_events is how many of calls do NOT carry metering_confidence = 'complete' — the TOKEN counts' companion, exactly as unknown_cost_events is the COST's.
A NON-ZERO VALUE MAKES ALL FOUR TOKEN SUMS A LOWER BOUND
Migration 203's COMMENT ON COLUMN llm_usage_events.input_tokens is normative and it is addressed to this surface:
"A 0 here is only meaningful when metering_confidence = 'complete'. On a 'blind' row it means NOT OBSERVED, not zero. ANY CONSUMER THAT SUMS TOKENS MUST CARRY A COMPANION COUNT OF NON-COMPLETE ROWS, or gate on metering_confidence."
pricing.Price implements the same rule as its very first check — if u.Confidence != ConfidenceComplete { return unknown(metering_blind) } — so a row this field counts is one the pricer treats as blind.
WHY IT IS NOT A COUNT OF metering_confidence IS NULL
NULL means NOT ASSESSED and is the correct value for every row predating migration 203 and for the two worker-side producers. But 'partial' and 'blind' are equally not-totals — 'partial' is defined as "some classes were observed, others were not; the observable classes are a LOWER BOUND" — and the AI Gateway's emitUsage writes 'partial' explicitly. A count of NULLs alone is therefore correct on today's corpus (which holds no assessed rows at all) and wrong on the first row a metering-aware writer produces: wrong by exactly the one value nobody listed. The predicate is IS DISTINCT FROM 'complete', which is the predicate Price uses.
WHY IT IS INDEPENDENT OF unknown_cost_events, AND MUST STAY SO
A row can be UNMETERED AND PRICED. The worker-side producers write a real, worker-reported cost_usd and assess no confidence, so they are priced with unassessed tokens — measured on the live corpus at review time: 193 rows, 193 unassessed, ZERO cost-unknown. Cost honesty read "perfect" while every token total was a lower bound. Do not derive either count from the other and do not add an invariant relating them. |
| total_ms_sum | int64 | optional | total_ms_sum is the summed wall-clock duration, in milliseconds, of the calls in this bucket that carry an OBSERVED total_ms (> 0).
ABSENT means "no call in this bucket recorded a duration". Exactly as with cost_usd, it MUST NOT be rendered as zero: a bucket of unmeasured calls took an unknown amount of time, not none. The absent-versus-zero distinction is carried in the TYPE — an optional — for the same reason it is on cost_usd, and it is the one input on which an honest and a COALESCE-ing implementation disagree. |
| total_ms_measured_events | int64 | | total_ms_measured_events is how many of calls contributed to total_ms_sum — the denominator any mean derived from it was computed over. |
| total_ms_unmeasured_events | int64 | | total_ms_unmeasured_events is how many of calls did NOT: NULL, the writer's 0 sentinel, or a negative value. It is total_ms_sum's companion exactly as unknown_cost_events is cost_usd's, and a non-zero value means the sum is a LOWER BOUND on the bucket's real wall clock.
It is a SEPARATE aggregate rather than calls - measured: the server asserts measured + unmeasured == calls on every row and on the totals before answering, and a subtraction would make that identity a tautology instead of a check. |
| ttft_ms_sum | int64 | optional | ttft_ms_sum is the summed time-to-first-token, in milliseconds, over the calls carrying an OBSERVED ttft_ms (> 0). Same absent-versus-zero rule as total_ms_sum.
NOTE THE POPULATION IS NARROWER THAN total_ms's. ttft_ms is meaningless for a NON-STREAMING call and is 0 there (migration 227's COMMENT ON COLUMN), so a bucket of non-streaming calls reports this ABSENT with ttft_ms_unmeasured_events == calls. That is the correct answer, and it is why the two metrics carry their own denominators rather than sharing one. |
| ttft_ms_measured_events | int64 | | ttft_ms_measured_events is how many of calls contributed to ttft_ms_sum. |
| ttft_ms_unmeasured_events | int64 | | ttft_ms_unmeasured_events is how many did not — including every non-streaming call. |
| no_termination_recorded_events | int64 | | no_termination_recorded_events counts the calls whose termination_reason IS NULL.
IT IS NOT "RAN TO COMPLETION", AND THE NAME SAYS SO ON PURPOSE
Migration 219's termination_only_on_allow CHECK permits a reason ONLY beside governance_outcome = 'allow' — never beside a denial and never beside an UNASSESSED row. So for the whole pre-migration-203 corpus and for both worker-side producers, a cut COULD NOT HAVE BEEN RECORDED even if one happened, and every one of those rows is counted here.
Reading this as "completed successfully" is the same error as reading a NULL governance_outcome as allow: it reports a clean run that was never observed. A consumer that needs "ran to completion" needs the governance axis too — that is GovernancePanel's surface (#3023) and deliberately not this one. |
| terminated_duration_cap_events | int64 | | terminated_duration_cap_events counts termination_reason = 'duration_cap' — the drain-safety ceiling cutting a generation that WAS producing frames.
LLD #2801 §8.5: this is CORRECT BEHAVIOUR, not an incident. A single-shot generation longer than the cap is out of Model Gateway scope. Do not render it as an error rate. |
| terminated_client_gone_events | int64 | | terminated_client_gone_events counts termination_reason = 'client_gone' — the caller hanging up, or its context being cancelled, before the stream finished. The cost was still incurred, which is why it appears on a spend surface at all. |
| terminated_upstream_stalled_events | int64 | | terminated_upstream_stalled_events counts termination_reason = 'upstream_stalled' — the ceiling firing having DELIVERED ZERO FRAMES to the client. Read that literally: the discriminator is deliveries observed on the near side, not a claim about the far side (review N2 on PR #2862). |
| terminated_unrecognised_events | int64 | | terminated_unrecognised_events counts calls carrying a NON-NULL termination_reason that is none of the three codes above.
IT IS ZERO ON A COHERENT DATABASE, AND IT IS STILL SENT
termination_reason_domain makes a fourth code unstorable, so with that CHECK intact this is always 0. It exists because the alternative is worse in exactly one direction: without it the five counters would not partition calls, so a code that DID get stored — through a constraint dropped for a backfill, a restore, or a fourth code added to the domain without this file — would be silently absent from every bucket while calls still counted it. That is the drop-a-population defect, and the server's partition check turns it into a refusal instead.
It is also the mechanical rendering of "an unrecognised code renders as unknown rather than being coerced into a nicer bucket": such a code lands here, NOT in no_termination_recorded_events. |
| governance_assessed_allow_events | int64 | | governance_assessed_allow_events counts governance_outcome = 'allow' — the gate chain RAN and permitted the call.
IT IS NOT "the calls that were fine". Read it only against governance_unassessed_events: allow/(allow+deny) is a rate over the assessed population, and the assessed population may be a small and unrepresentative slice of calls. |
| governance_assessed_deny_events | int64 | | governance_assessed_deny_events counts governance_outcome = 'deny' — a gate refused the call. Migration 207's gate_iff_not_allow makes deciding_gate non-NULL on exactly these rows, which is why governance_gate_named_events below must equal this number and the server refuses when it does not. |
| governance_assessed_unrecognised_events | int64 | | governance_assessed_unrecognised_events counts a NON-NULL governance_outcome that is neither allow nor deny.
IT IS ZERO ON A COHERENT DATABASE, AND IT IS STILL SENT
Migration 203's domain CHECK makes a third value unstorable, so with that CHECK intact this is always 0. It exists for the reason terminated_unrecognised_events does: without it the four counters would not partition calls, so a value that DID get stored — through a constraint dropped for a backfill, a restore, or a third value added to the domain without this surface — would be silently absent from every bucket while calls still counted it.
Note where such a value does NOT land: not in governance_assessed_allow_events, and not in governance_unassessed_events either. Folding it into the latter would be the same lie in the other direction — an assessment DID happen and its verdict is unreadable, which is a different fact from no assessment. |
| governance_unassessed_events | int64 | | governance_unassessed_events counts governance_outcome IS NULL — NO gate chain evaluated this call.
THIS IS THE HONEST DENOMINATOR, AND ON TODAY'S LEDGER IT IS MOST OF IT
It is a first-class field rather than something a client derives, for the same reason unmetered_events is: the value cannot be recovered from the others by anyone who does not already know the rule, and the client that gets the rule wrong gets it wrong in the flattering direction.
It is NOT an error, a gap, or a backlog. It is the correct record of a call that was never governed — every pre-203 row, and every call from the two worker-side producers, which do not pass through the gate chain by design. A panel that renders this as a failure is as wrong as one that renders it as an allow.
The name says unassessed and not ungoverned deliberately: the column records whether a VERDICT was written, which is a narrower claim than whether any governance existed anywhere in the path. |
| governance_gate_named_events | int64 | | governance_gate_named_events counts deciding_gate IS NOT NULL — calls whose refusal names the gate that refused it.
IT IS A CROSS-ORACLE, NOT A SECOND SPELLING OF THE DENY COUNT
Migration 207's gate_iff_not_allow is (governance_outcome IS NOT DISTINCT FROM 'deny') = (deciding_gate IS NOT NULL) — null-safe, so it cannot be satisfied vacuously by an unassessed row. With it intact this field EQUALS governance_assessed_deny_events exactly, and the server asserts that and REFUSES when it does not hold.
The two numbers reach that answer through DIFFERENT COLUMNS, which is what makes the equality worth checking: a divergence is a denial that names no gate (unauditable) or a gate named beside a call nobody refused (the vacuous arm 207 closed). Both are database-invariant violations and both render as an ordinary-looking panel, so the read surface is the only place either can be caught.
THE GATE NAMES THEMSELVES ARE NOT BROKEN OUT HERE, and #3047's acceptance asked for per-gate counters over a closed vocabulary. THAT ASK WAS WITHDRAWN AND THE COUNTERS ARE REJECTED FOR v1 — not deferred, and NOT a v1.1 backlog item. The architect re-derived the measurement below by executing the code and withdrew the ask on the evidence (#3050 review). Filing "add per-gate counters" as a follow-up would re-open a decision that is settled; the reason is a missing PROPERTY, not a missing sprint.
The property: deciding_gate HAS NO CLOSED VOCABULARY TO COUNT OVER. Unlike governance_outcome (migration 203) and termination_reason (migration 219), it carries NO domain CHECK, and its value set is the UNION of two producers' refusal vocabularies:
internal/modelgateway governanceDenialReasons() 47 internal/aigateway the Gate* consts 5 (2 shared) ── UNION 50
Measured 2026-09-06; re-derive rather than trusting the digits — the first draft of this comment said "~50 stable strings" as an estimate and happened to be right, which is not a method. governanceDenialReasons() is itself DERIVED from refusalContract (every spec whose Ledger is LedgerDeny), so it grows silently with every new refusal.
50 open-ended, silently-growing strings is not the shape terminated_*_events uses — that is 3 members fixed by a CHECK. Fifty per-gate FIELDS would be a wire surface that has to be edited on every new refusal, and an ..._unrecognised_... remainder over a vocabulary with no CHECK behind it would be COUNTING DRIFT RATHER THAN DEFECTS — a counter whose value is noise.
It is the same reasoning that ruled out a USAGE_GRAIN_GATE when #3047 was scoped: deciding_gate exists only on non-allow rows, so a gate breakdown silently defines its own denominator.
THE PREREQUISITE IS A CLOSED VOCABULARY, NOT MORE FIELDS. If a per-gate breakdown is ever wanted, the correct FIRST move is a domain CHECK or a lookup table for deciding_gate — after which the breakdown is a five-line derivation off one declaration, exactly as governanceOutcomes is here. Asking for the fields before that property exists is asking for the defect.
What ships instead is the cross-oracle above, which needs NO vocabulary at all and is a genuine second reading rather than a second spelling of the deny count. See #3047 and the #3050 review. |
| governance_gate_unnamed_events | int64 | | governance_gate_unnamed_events counts deciding_gate IS NULL — allowed, unassessed, and (only on an incoherent database) a denial that names no gate.
A SEPARATE aggregate rather than calls - gate_named, for the reason every other complement on this message is: the server asserts gate_named + gate_unnamed == calls, and a subtraction would make that a tautology instead of a check.
IT IS AMBIGUOUS ON ITS OWN and the name does not pretend otherwise — migration 227's COMMENT ON COLUMN says deciding_gate "is readable only TOGETHER with governance_outcome". Read it beside the four outcome counters or not at all. |
UsageFilter
UsageFilter is the selection shared by both RPCs.
It narrows WITHIN the caller's org. There is no org field: the org is the authenticated caller's and cannot be named by the body.
| Field | Type | Label | Description |
|---|---|---|---|
| grain | UsageGrain | grain is the dimension to group by. Required; UNSPECIFIED is refused. | |
| window | UsageWindow | window is the half-open time range. Required. | |
| team_id | string | team_id narrows to one llm_usage_events.team_id. Empty does not narrow. |
This is a FILTER and is orthogonal to the grain: filtering to one team at model grain is "what did this team spend, by model". |
| agent_id | string | | agent_id narrows to one agent, matched against the UUID agent_id and the opaque agent_id_text alike. Empty does not narrow. |
| model | string | | model narrows to one model id, matched exactly. Empty does not narrow. |
UsageRow
UsageRow is one bucket at the requested grain.
| Field | Type | Label | Description |
|---|---|---|---|
| bucket_key | string | bucket_key is the requested grain's value, and it is OPAQUE TEXT whose meaning is the grain's. |
Each UsageGrain member documents its own key — a UUID, an opaque runtime id, a model name, a YYYY-MM-DD day. THIS COMMENT DELIBERATELY DOES NOT ENUMERATE THEM: a hand-written list inside prose is the one place drift is invisible to every descriptor-derived guard, and it was accurate here until the first grain that was not an id. Read the enum, not this sentence, and do not parse the key against a per-grain shape you inferred from a sample.
EMPTY IS A REAL BUCKET, not a missing one. Several grains have a population whose source column is legitimately NULL — no team, no agent (migration 203 made agent_id nullable precisely so a human/CI caller could be recorded), no endpoint (a worker-side producer rather than the Model Gateway) — and those events are collected under the empty key rather than dropped. A grouping that drops rows makes the page's counts disagree with the totals, which is the defect the server's own partition check exists to catch. A grain with no such population declares that too, and carries no COALESCE. |
| aggregate | UsageAggregate | | aggregate is the bucket's counted shape. |
UsageTotals
UsageTotals is the aggregate over the WHOLE filtered window, independent of pagination.
It is always present, including on an empty corpus — where every counter is
zero and cost_usd is ABSENT. An empty result that omitted the totals would
be indistinguishable from a page the client failed to read.
| Field | Type | Label | Description |
|---|---|---|---|
| buckets | int64 | buckets is how many distinct grain keys the filter matched across every page. It is the dimension a page-local count cannot supply, and it is what makes "the pages summed to the total" checkable by the client. | |
| aggregate | UsageAggregate | aggregate is the same counted shape as a row's, summed over every bucket. |
UsageWindow
UsageWindow is the half-open time range [since, until) the query covers.
BOTH BOUNDS ARE REQUIRED. There is deliberately no "all time" default and no implicit trailing window: an unbounded scan over an append-only ledger is a cost the caller should have to ask for, and a defaulted window silently changes the answer as the corpus grows.
| Field | Type | Label | Description |
|---|---|---|---|
| since | google.protobuf.Timestamp | since is the INCLUSIVE lower bound. Required. | |
| until | google.protobuf.Timestamp | until is the EXCLUSIVE upper bound. Required, and must be strictly after since. Half-open so two adjacent windows partition the ledger exactly once — a closed upper bound double-counts the boundary event. |
UsageGrain
UsageGrain is the dimension the ledger is grouped by. It decides the GROUP BY key, so the set is closed — see the header.
| Name | Number | Description |
|---|---|---|
| USAGE_GRAIN_UNSPECIFIED | 0 | USAGE_GRAIN_UNSPECIFIED is REFUSED with InvalidArgument. It is not a default: a grain that was not asked for answers a different question. |
| USAGE_GRAIN_ORG | 1 | USAGE_GRAIN_ORG collapses the whole window into a single bucket keyed by the caller's org. It is the "what did we spend" number. |
| USAGE_GRAIN_TEAM | 2 | USAGE_GRAIN_TEAM groups by llm_usage_events.team_id. Events carrying no team land in the unattributed bucket (empty bucket_key) rather than being dropped — see UsageRow.bucket_key. |
| USAGE_GRAIN_AGENT | 3 | USAGE_GRAIN_AGENT groups by the agent's id, preferring the UUID agent_id and falling back to the opaque agent_id_text that migration 203 added for callers whose identity is not a UUID. A human/CI caller is neither, and lands in the unattributed bucket. |
| USAGE_GRAIN_MODEL | 4 | USAGE_GRAIN_MODEL groups by llm_usage_events.model, which is NOT NULL — so this grain has no unattributed bucket. |
| USAGE_GRAIN_ENDPOINT | 5 | USAGE_GRAIN_ENDPOINT groups by llm_usage_events.endpoint_id, the LLM endpoint a call was routed through. |
IT HAS AN UNATTRIBUTED BUCKET, AND ON TODAY'S CORPUS IT IS THE WHOLE OF IT. Migration 203's COMMENT ON COLUMN is explicit that endpoint_id is "NOT NULL iff the Model Gateway produced this event", so every row written by the two worker-side producers carries no endpoint. Those rows land in the empty-bucket_key bucket — a REAL bucket meaning "not produced by the Model Gateway", counted like any other — rather than being dropped. A breakdown that silently omitted them would report the gateway's spend as the org's total. |
| USAGE_GRAIN_DAY | 6 | USAGE_GRAIN_DAY groups by the UTC calendar day of llm_usage_events.created_at, which is TIMESTAMPTZ NOT NULL (migration 018) — so this grain has NO unattributed bucket, and the key expression carries no COALESCE that would imply one exists.
bucket_key is the day rendered as YYYY-MM-DD. THE ZONE IS PINNED TO UTC rather than left to the session's TimeZone: a day boundary that moves with the server's locale makes the same window return different buckets from two connections, and the difference shows up as spend migrating between two adjacent days rather than as an error. YYYY-MM-DD also sorts lexicographically in chronological order, which is what lets the shared ascending bucket_key keyset cursor page a trend chart correctly. |
| USAGE_GRAIN_SESSION | 7 | USAGE_GRAIN_SESSION groups by llm_usage_events.session_id, the agent session a call was made inside. It is AgentEconomicsTable's grain (MG-FE.4b, #3023): "what did this conversation cost".
IT HAS AN UNATTRIBUTED BUCKET AND THE BUCKET IS THE MAJORITY POPULATION. session_id is UUID REFERENCES agent_sessions(id) and NULLABLE (migration 018), and it is written only by a producer that HAS a session: a one-shot Model Gateway call, a CI caller and a human caller all record NULL. Those rows land in the empty-bucket_key bucket — a REAL bucket meaning "not made inside a session", counted like any other — rather than being dropped.
Dropping them would be the same defect as at ENDPOINT grain and it reads the same way: the sessioned share of the ledger presented as the org's total, i.e. a SMALLER BILL rather than an error. The proof is a count across buckets against the totals, not a presence assertion — see TestModelGatewayUsage3027_SessionEmptyBucketIsCountedNotDropped. |
| USAGE_GRAIN_TRIGGERING_TEAM | 8 | USAGE_GRAIN_TRIGGERING_TEAM groups by llm_usage_events.triggering_team_id (migration 248): WHO CAUSED a call, as opposed to USAGE_GRAIN_TEAM's team_id, which is the team the CREDENTIAL is scoped to and therefore WHO PAID (ADR-0035 decision D3, MG-CTX.10).
THESE NUMBERS ARE NOT FUNDING AND NOT CHARGEBACK, and that is a constraint on the requirement rather than a caveat on the implementation. An ORG-GRAIN capability credential — the platform-minted, caller-less background work of ADR-0035 D1 (embedding, compaction, extraction, moderation, rag_generation) — is scoped to no team at all, so nothing here says who should be billed. It says which team's activity CAUSED the work. The funding half of the same row lives on the ledger as credential_class / credential_runtime (migration 225) and is deliberately not projected into the view or into any grain. Two honest columns are worth more than one column with a contested name.
IT HAS AN UNATTRIBUTED BUCKET AND THE BUCKET IS THE MAJORITY POPULATION — on today's corpus it is ALL of it. triggering_team_id is written only by a capability caller that was handed a triggering unit; every member-API-key call, every agent call, every worker-side row and every scheduled capability job (a nightly re-embedding has no triggering team) records NULL. Those rows land in the empty-bucket_key bucket — a REAL bucket meaning "no triggering team was recorded", counted like any other — and NOT dropped.
The empty bucket here is therefore LARGE AND CORRECT, and it must never be read as "spend nobody caused": it is spend whose causation this ledger cannot answer for. Summing only the named buckets and presenting the result as the org's total reports a smaller number than the truth, which is the same honesty rule cost_unknown carries beside cost_usd (MG-1.2). |
AIGatewayUsageService
AIGatewayUsageService is the tenant-facing READ surface over
llm_usage_effective_cost — what was spent, by whom, on what.
Every RPC is scoped to the caller's org, taken from the authenticated caller
and NEVER from the request body, and every read runs inside the request's
RLS-bound scope transaction so migration 018's
llm_usage_events_org_isolation policy is the isolation mechanism. The view
is security_invoker = true, so the base tables' policies are evaluated as
the querying role rather than as the view's owner.
READ-ONLY BY CONSTRUCTION. There is no RPC on this service that writes
anything, and in particular none that writes llm_rate_cards — see the
#2876 note in this file's header.
| Method Name | Request Type | Response Type | Description |
|---|---|---|---|
| QueryUsage | QueryUsageRequest | QueryUsageResponse | QueryUsage returns one page of aggregated usage at the requested grain, plus the totals for the WHOLE filtered window (not for the page). |
Requires L3, the same floor as every other Model Gateway read surface. | | ExportUsage | ExportUsageRequest | ExportUsageResponse | ExportUsage returns the same page rendered as CSV, plus the same totals as a structured message.
The totals ride ALONGSIDE the CSV rather than as a trailing row inside it: a totals row in the body is indistinguishable from a bucket to any consumer that sums the column, which is how an export comes to double-count itself. Requires L3. | | QueryLatency | QueryLatencyRequest | QueryLatencyResponse | QueryLatency returns one page of LATENCY DISTRIBUTIONS at the requested grain, plus the distribution for the WHOLE filtered window.
It exists as a separate RPC — rather than as extra fields on QueryUsage — for the reason spelled out at length on LatencyDistribution: percentiles do not compose by addition, and UsageAggregate is the SUMMED shape. Read that comment before proposing a merge of the two.
Same UsageFilter, same UsageWindow semantics, same L3 floor, same keyset pagination and the same page tokens as QueryUsage: it groups the SAME rows by the SAME key expression, so a token minted by either RPC positions the other identically. Only the statistic differs.
READ-ONLY, like the two above. It writes nothing and in particular writes no rate card — see the #2876 note in this file's header. |
Scalar Value Types
| .proto Type | Notes | C++ | Java | Python | Go | C# | PHP | Ruby |
|---|---|---|---|---|---|---|---|---|
| double | double | double | float | float64 | double | float | Float | |
| float | float | float | float | float32 | float | float | Float | |
| int32 | Uses variable-length encoding. Inefficient for encoding negative numbers – if your field is likely to have negative values, use sint32 instead. | int32 | int | int | int32 | int | integer | Bignum or Fixnum (as required) |
| int64 | Uses variable-length encoding. Inefficient for encoding negative numbers – if your field is likely to have negative values, use sint64 instead. | int64 | long | int/long | int64 | long | integer/string | Bignum |
| uint32 | Uses variable-length encoding. | uint32 | int | int/long | uint32 | uint | integer | Bignum or Fixnum (as required) |
| uint64 | Uses variable-length encoding. | uint64 | long | int/long | uint64 | ulong | integer/string | Bignum or Fixnum (as required) |
| sint32 | Uses variable-length encoding. Signed int value. These more efficiently encode negative numbers than regular int32s. | int32 | int | int | int32 | int | integer | Bignum or Fixnum (as required) |
| sint64 | Uses variable-length encoding. Signed int value. These more efficiently encode negative numbers than regular int64s. | int64 | long | int/long | int64 | long | integer/string | Bignum |
| fixed32 | Always four bytes. More efficient than uint32 if values are often greater than 2^28. | uint32 | int | int | uint32 | uint | integer | Bignum or Fixnum (as required) |
| fixed64 | Always eight bytes. More efficient than uint64 if values are often greater than 2^56. | uint64 | long | int/long | uint64 | ulong | integer/string | Bignum |
| sfixed32 | Always four bytes. | int32 | int | int | int32 | int | integer | Bignum or Fixnum (as required) |
| sfixed64 | Always eight bytes. | int64 | long | int/long | int64 | long | integer/string | Bignum |
| bool | bool | boolean | boolean | bool | bool | boolean | TrueClass/FalseClass | |
| string | A string must always contain UTF-8 encoded or 7-bit ASCII text. | string | String | str/unicode | string | string | string | String (UTF-8) |
| bytes | May contain any arbitrary sequence of bytes. | string | ByteString | str | []byte | ByteString | string | String (ASCII-8BIT) |