Skip to main content

Completion-path provider posture, per deployment shape

Status: normative. PRD #996 FR-16 · LLD #2714 §5 T10 · ADR-0031 D4.

This page states, per deployment shape, which LLM providers the completion path registers and what that means for egress. It exists because one code comment asserted a posture the mechanism never encoded, and because the GTM then inverted the shape that comment assumed.

Machine-readable source of truth: internal/testsupport/gatewayshape. Every shape below is a row in gatewayshape.Shapes, and TestRunbookStatesEveryShapesPosture fails if this page and that table disagree. Edit the table, then edit this page.


The claim this page retires​

internal/aigateway/provider/local.go used to say:

PRODUCTION POSTURE (fail closed): Local is registered ONLY when AI_GATEWAY_LOCAL_BASE_URL is set (dev stack). It is absent in prod/beta, so the gateway there serves cloud BYOK providers exclusively.

The first sentence is the mechanism and is true. The second is a policy claim the code never made: nothing on the registration path reads ENVIRONMENT, and nothing reads a deployment-shape flag. "It is absent in prod" described how we happened to fill in a values file, in the grammar of a guarantee.

For a hosted product that was harmless. For on-prem it is backwards — there, "fail closed" means local only, not cloud only — and an operator reading it would conclude the code was about to fight them.

No behaviour changes as a result of writing this down. The mechanism was already shape-agnostic and already correct. What was missing was the statement.


The rule​

The local provider is registered exactly when AI_GATEWAY_LOCAL_BASE_URL names an endpoint — in every deployment shape, in every environment, with no override flag.

That single input is deliberate, and it is ADR-0031 D4 applied to this seam. D4 names AI_GATEWAY_LOCAL_BASE_URL in its own trust table as the archetype of a trusted, process-configuration URL, in the same class as DATABASE_URL: whoever can set it can already set the database URL, so there is no privilege boundary left to defend and no environment check worth making.

D4 is also explicit about the alternative: "keying on shape merely requires a shape flag, and a flag that defaults wrong reproduces the ENVIRONMENT == "production" bug in a new costume."


Per shape​

shapesets AI_GATEWAY_LOCAL_BASE_URLlocal providerhosted providers (anthropic, openai)default egress
devyes — the in-stack Ollama, http://ollama:11434/v1registered, keyless, priority 50registered, inert without a BYOK keynone
on-premyes — the operator's own in-perimeter inference serverregistered, the same code path as devregistered, inert without a BYOK keynone
hosted (SaaS)nonot registered; a local-family model id resolves to no provider and is rejectedregistered; each tenant's own BYOK key selects the vendoronly what a tenant configured

dev​

dev sets AI_GATEWAY_LOCAL_BASE_URL to the in-stack Ollama, so the keyless local provider is registered and local-family models are served inside the compose network.

It used to be set in docker-compose.dev.yml under the ai-gateway service. T18 (#2920) deleted that binary, and with it the only caller of provider.NewRegistry — so nothing in the dev stack reads AI_GATEWAY_LOCAL_BASE_URL any more and the compose entry went with the service. internal/aigateway/provider/{local,posture}.go survive (#3202) and still implement the rule; they simply have no composition root today.

The live mirror of this posture is the worker's WORKER_OLLAMA_BASE_URL, which is untouched and is where a reader should look for behaviour. The Go-side statement below is a decision that is still written down and still guarded — it is not currently executed by any running process.

on-prem​

on-prem sets AI_GATEWAY_LOCAL_BASE_URL to the operator's own in-perimeter inference server and gets the same code path as dev; a private address is the expected production configuration here and needs no override flag.

# Any OpenAI-compatible /v1 endpoint: Ollama, vLLM, TEI, llama.cpp server, a
# vendor appliance. An RFC1918 address, a Kubernetes service name and a
# private DNS suffix are all fine, and none of them require a flag.
AI_GATEWAY_LOCAL_BASE_URL=http://inference.internal:8000/v1

The models this serves are the families in provider.LocalModelPrefixes (gemma, qwen, llama, mistral, phi, deepseek), matched case-insensitively by prefix, so both a bare family name and a concrete tag (qwen2.5:3b) route locally. None of those prefixes overlap a hosted vendor's model ids, so a local model can never be misrouted to a cloud provider.

PRD #996 G2 is the property to preserve here: dev and on-prem are the same code path. The gateway cannot tell them apart, and must not learn how.

hosted (SaaS)​

hosted (SaaS) leaves AI_GATEWAY_LOCAL_BASE_URL unset, so no local provider is registered and a local-family model id resolves to no provider and is rejected, rather than being served by a cloud provider that happens to be registered.

The rejection is the point: "falls through to OpenAI" and "is refused" look the same if you only check that the local provider is absent.


What this posture does not claim​

It does not claim the gateway cannot egress. The hosted anthropic and openai providers are registered unconditionally, in every shape, including on-prem, each pointed at its vendor's public API (api.anthropic.com, api.openai.com). Nothing removes them.

What keeps a default install from reaching either is not their absence:

  1. Neither advertises the provider.Keyless capability, so handler.prepare resolves a BYOK key from the vault before either provider is dialled.
  2. With no key configured for the organisation, the request ends in 401 no_provider_key — the vendor is never contacted.
  3. Ahead of that, the per-team model allowlist and the budget pre-check both run and both fail closed.

So a default install has no egress it did not configure, and the mechanism that delivers that is the credential requirement, not provider absence. An on-prem operator who wants a hard guarantee should not rely on this page: put the egress control at the network, where it belongs.

This is pinned by TestOnlyTheLocalProviderIsKeyless (internal/aigateway/provider/posture_test.go), which walks the provider registry and fails if any hosted provider becomes keyless.

A weakening worth recording rather than glossing. Until T18 (#2920) the pin was TestNoShapeShipsAProviderThatCanEgressWithoutAnOperatorCredential in cmd/ai-gateway/posture_test.go, and it drove the production composition root in each shape — it asserted the property of the thing that actually booted. T18 deleted that composition root, so no test can drive it any more. What survives asserts the same property of the registry, one layer in. The gap between them is a wiring mistake in a composition root that no longer exists, so nothing is currently unguarded; but if provider ever regains a composition root (see #3202), the composition-root-level assertion should come back with it.


Two seams, two trust classes — and why only one gets a tripwire​

The completion path has both rows of ADR-0031 D4's trust table, and treats each according to its provenance. This is the reason the embedding seam's "sourced only from process configuration" tripwire (test/lint/embedding_url_provenance_test.go, LLD #2714 T11) is not copied wholesale onto the completion path: applied to the whole completion path it would be false, and a tripwire that must be suppressed is worse than none.

seamprovenancetrustguard
AI_GATEWAY_LOCAL_BASE_URL → provider.NewLocal (internal/aigateway/provider; its cmd/ai-gateway composition root was deleted by #2920)process configuration (row 1)trustednone, by design. A narrow inventory tripwire (test/lint/completion_path_posture_test.go) reddens if a DB-sourced or tenant-scoped path to this value ever appears
model_endpoints.base_url → cmd/model-gatewaytenant-supplied and stored (row 2)untrusted inputthe full SSRF fence: internal/modelgateway/registry.BaseURLValidator, re-validated on every edit, L5 on every status transition

Verifying the posture on a running install​

The gateway logs the resolved posture once at startup:

level=INFO msg="completion-path provider posture resolved"
local_provider_registered=true
local_endpoint=http://inference.internal:8000/v1
local_endpoint_class=private
providers="[local anthropic openai]"
local_endpoint_var=AI_GATEWAY_LOCAL_BASE_URL

local_endpoint_class is ADR-0030's bounded vocabulary (none / private / google / external / unknown), so it stays cardinality-safe even though the hostname is operator-supplied. A local endpoint that reads external means the "local" provider is pointed at a third-party host — permitted (row 1 trusts the operator) but worth a second look.

The log line deliberately does not name a deployment shape. The process does not know which shape it is running in, and a line that guessed would be the same class of claim this page exists to retire.


  • ADR-0031 §D4 — SSRF guard keys on URL provenance, not environment or shape.
  • PRD #996 FR-16, G2, G4.
  • internal/aigateway/provider/posture.go — the statement at the decision point.
  • internal/testsupport/gatewayshape — the machine-readable table.
  • docs/runbooks/beta-vertex-embeddings.md — the embedding path's equivalent.