ADR-0021: Registry & Gateway Architecture (MCP + LLM)
- Status: Accepted (founder approval 2026-07-06,
upsquad-ai/upsquad-core#1682) - Date: 2026-07-06
- Decision owners: Founder (sign-off), Principal Architect (author), Backend SME (Context Engine / gateway), DevOps (GKE LOE), Product Manager (PRD #1674)
- Related:
- PRD
upsquad-ai/upsquad-core#1674 - ADR-0020 — dual-runtime worker types (
#1627) upsquad-ai/upsquad-core#1647(+ its W2 dependency for grants reaching agents)upsquad-ai/upsquad-core#1642/#1643— Org v2.4- LiteLLM dev facade PR
#1632
- PRD
Status
Accepted — founder approval 2026-07-06 (via GitHub comment on #1682).
Context
Founder requirements driving this ADR:
- MCPs must be registry-mandatory (no ad-hoc upstreams).
- Each MCP gets a stable per-MCP proxy URL.
- Team binding with team-owner/Manager approval; cascade via Org v2.4 resolver + the
core#1647snapshot work. - User-composable group proxies that aggregate multiple MCPs without tool-name conflicts.
- Per-tool deterministic guardrails set at team level (e.g.
args.amount <= N, response filters). - The same architecture pattern applied to LLM models (model registry + gateway), not just MCP tools.
This ADR captures ONE architecture pattern with TWO instances: MCP (tools) registry+gateway and LLM (models) registry+gateway.
Basis: adversarially verified research (deep-research workflow, 9 findings, 0 refuted), code-grounded architect passes over the current ContextForge-based mcp-proxy + Go registry, and DevOps/LOE assessments.
Decisions
1. Registry unification. proxy_upstreams becomes a projection of mcp_servers + mcp_server_units — a single source of truth. The registry gains probe-time tool-schema capture, versioning, and status = pending | approved | disabled. Schema modelled on the official MCP registry (namespaced names, ownership verification, designed for private-subregistry federation — verified 6-0).
2. Proxy technology. KEEP the ContextForge-based mcp-proxy + Go registry (ContextForge's two-tier auth verified 9-0 as the reference multi-tenant pattern). Envoy stays edge-only. NO rebuild. Envoy AI Gateway verified 12-0 as the strongest k8s-native alternative (MCPRoute aggregation, session multiplexing, OAuth termination, native per-tool CEL) — a formal RE-EVALUATION GATE at GKE migration.
3. URL scheme. App-layer path routing: /mcp/{registry-id} and /mcp/group/{id}; the edge keeps ONE coarse route. Each URL serves per-path RFC 9728 protected-resource metadata (spec-mandated, verified 6-0).
4. Auth. Gateway terminates OAuth; token passthrough is NORMATIVELY BANNED (verified 9-0, spec verbatim). Per-upstream tokens via RFC 8693 token exchange; RFC 8707 audience binding mandatory; confused-deputy prevention (per-client consent + client registry) required for dynamic client registration (verified 6-0).
5. Groups. Register-time conflict rejection for authored groups + automatic {server}__{tool} prefixing as safety net (industry consensus verified 15-0). Group session-map = one client session → N upstream sessions. v1: streamable-HTTP upstreams only (SSE federation unsupported industry-wide); reject callback-requiring upstreams.
6. Per-tool guardrails. Policy store in Postgres keyed (team, server, tool); ARGUMENT rules as CEL expressions + RESPONSE filters (DLP/regex/redaction); evaluated in our proxy plugin chain; policies EMIT portable CEL so they can compile onto Envoy AI Gateway post-GKE. Team-level semantics ride the Org v2.4 guardrail axis (unconditional push down home chain).
7. State discipline. Group session-map + snapshot caches in REDIS, never in-process (multi-replica GKE requirement).
8. Approval. MCP↔team attach requires team Manager approval (org_units.manager_member_id, Org v2.4 authority); binding is inert until approved.
9. LLM instance of the pattern. Model catalog + BYOK vault + team_model_allowlists (exists) + internal/aigateway + Anthropic-format facade (replaces dev LiteLLM; ties to PRD #1674); same team binding / approval / guardrail semantics.
Amendment, 2026-09-25 (core#3601; founder direction, architect's ruling
#3601 issuecomment-5828940160§1/§3). The LLM instance's approval axis is the ENDPOINT BINDING; the model set is a narrowing ATTRIBUTE of it.Decision 9 gave the LLM instance a model-level approval axis mirroring the MCP one:
team_model_allowlistswith its own pending/approved/denied lifecycle and its own Manager inbox. That mirror is withdrawn. The two grains were not analogous — an MCP server is a resource a team is granted, while a model is something an already-granted endpoint either serves or does not — and keeping a separate approval axis for the second produced a control that could name models the team could not reach and exclude the ones it could. Measured on beta-dev: the Engineering team held two PENDING entries (gpt-4o,Claude Sonnet 4), neither served by any endpoint it was bound to, and approving either would have made gate 7c deny the only model its endpoint served (core#31465813201959).What replaces it:
llm_endpoint_units.enabled_models(migration 262), a narrow-only JSONB attribute of the binding row. Binding approval (L4 and the unit's Manager) approves the ceiling — every model that endpoint serves — and the narrowing can only remove from it, so it needs no second approver. NULL and[]both mean no narrowing; a list equal to the endpoint's full served set canonicalises to NULL. Gate 7c reads it off the binding row it already holds, so the model axis costs zero extra reads and cannot disagree with the endpoint axis by construction.Decision 8 is UNCHANGED and is the part that still holds: attachment requires team-Manager approval, and a binding is inert until approved. What changed is that the LLM instance now has ONE approval axis instead of two, which is what decision 8 describes for MCP.
team_model_allowlistsis retired as a policy source and is dropped in a later migration.
10. No service mesh now. Ambient mesh + sidecar + SPIFFE re-evaluated at GKE (verified: zero industry evidence of established practice on mesh L7 MCP awareness / per-agent sidecars / SPIFFE propagation — nothing to adopt yet).
GKE Transition
Sequencing decision = COMPOSE-FIRST. Rework analysis: throwaway is 3-5% (~half a day) — a few Envoy routes / compose defs. GKE-first would block the entire build 3-4 weeks and collide with the live Org v2.4 migrations on the same tables.
GKE migration facts (DevOps LOE): foundation ~80% exists (VPC / WIF / Artifact Registry / tunnel live; CloudSQL stopped with data; some kustomize manifests + 1 ArgoCD app). GkeStack is an EMPTY STUB.
Work items: fill GkeStack (M), missing manifests — ai-gateway / client / spa / pgbouncer / migrate-jobs (M), ArgoCD app-of-apps (M), CloudSQL reactivate (S), Memorystore recreate (S), Secrets → Secret Manager + ESO (M), Gateway API ingress + Cloudflare cutover (M), Managed Prometheus (M), CI rewire (M), cutover + rollback (M).
Total: 6-9 person-weeks, 3-4 weeks calendar, +$450-600/mo.
Autopilot constraints: no privileged daemonsets, no host networking, resource minimums, ollama impractical (prod = BYOK, drop).
What changes at GKE for THIS architecture: host Envoy → Gateway API; Envoy-AI-Gateway re-eval (decision 2 gate); ambient mesh for mTLS/identity; per-tenant worker pods make sidecars meaningful; SPIFFE identity revisit.
What stays identical: all app-layer routing, registries, policy store, session state (already Redis), plugins.
Alternatives Considered
- Adopt Envoy AI Gateway now — rejected: k8s-CRD-native, no compose deployment.
- Adopt agentgateway / MetaMCP — rejected: replaces the working governed proxy, loses our audit/idempotency plugins.
- GKE-first sequencing — rejected: blocks build 3-4 weeks, collides with Org v2.4 migrations.
- Build-nothing / static config — rejected: registry mandate + team cascade impossible.
Consequences
Positive
- Single source of truth (registry) removes
proxy_upstreamsdrift; versioned, probe-captured schemas. - One pattern reused for MCP and LLM → shared binding/approval/guardrail machinery.
- Portable CEL policies keep an exit path onto Envoy AI Gateway at GKE.
- Compose-first keeps the entire build unblocked; ~half a day of throwaway.
- Redis-first state is GKE-multi-replica-ready from day one.
Negative
- Continued dependency on the ContextForge upstream project (roadmap / maintenance risk).
- Group session-map complexity (1→N session fan-out, lifecycle, cleanup).
- Two proxy technologies coexist (our plugin proxy + edge Envoy) until the GKE re-eval gate.
Risks
- MCP auth spec churn (RFC 9728 / 8693 / 8707 / confused-deputy guidance still evolving) — may force follow-up work.
- SSE / callback upstreams unsupported in v1 → some real-world MCPs excluded until federation matures.
- Envoy AI Gateway re-eval could reverse decision 2, incurring migration cost of the plugin chain.