Skip to main content

Tenant install — the media.net GCP spoke

Issue core#2992, tracker #2991. ADR-0018 (cloud foundation), ADR-0017 (staged promotion), ADR-0034 (edge posture).

Read this first. This runbook describes two install shapes. They are not alternatives of equal standing:

ShapeStatusWhere
GCP spoke — private regional GKE Autopilot in the tenant's own projectPRIMARY. The media.net deliverable.§1–§9
Single-VM compose — one hardened docker compose profileSECONDARY. Generic on-prem artefact, kept for future prospects.§10

The compose profile was built first and was largely complete when the founder pivoted the target to a GCP project (2026-09-05). It is finished and shipped rather than deleted because a self-contained single-VM install has standalone value — but nothing about media.net's delivery runs on it.


1. What gets built, and what it costs to run​

Two artefacts, applied in order:

  1. infra/pulumi/upsquad-infra, stack mnet — the tenant's project: private VPC, Cloud NAT, regional private GKE Autopilot cluster, CloudSQL PostgreSQL 16 with regional HA and private-IP-only + enforced TLS, Memorystore Redis with a replica and in-transit encryption, a GCS bucket, Secret Manager, Workload Identity, a Cloudflare tunnel, and a budget.

  2. deployments/overlays/mnet — the workloads: Context Engine, Model Gateway, Agent Orchestrator, the memory MCP server, the knowledge MCP server, the approval scheduler, the client portal, the cloudflared connector, and the three database Jobs.

    Two of those were added by the architect's review of PR #3004. The approval scheduler because without it overdue governance approvals never expire — #2556 measured 71 pending approvals, every one past its expires_at, on a stack running without it. The knowledge MCP server because its base was always this PR's follow-through on #2993; note it is blocked from deploying until that binary gains TLS support for Redis (§9.4).

1.1 Security posture, stated as properties​

PropertyHow it is held
No inbound internet surfaceNo GCLB, no Service: LoadBalancer, no node external IPs, no Ingress object. Traffic arrives because the cloudflared pod dials out.
TLS at every off-node hopCloudflare edge TLS → tunnel (TLS) → in-cluster; CloudSQL sslMode: ENCRYPTED_ONLY; Memorystore SERVER_AUTHENTICATION + a static AUTH string; GCS native TLS.
Private control planeenablePrivateEndpoint: true, masterAuthorizedCidrs: []. kubectl needs a path into the VPC (§4.1).
No auth bypass reachableENVIRONMENT: production makes the gateway's bypass triple-gate structurally unreachable; DISABLE_AUTH/ICS_MODE/NEXT_PUBLIC_DISABLE_AUTH are false and asserted by preflight.
Tenant isolation live at runtimeServices connect as app_rw (NOSUPERUSER NOBYPASSRLS), so FORCE RLS applies. Smoke-tested (§7).
No service-account key filesVertex, GCS and the Drive connector authenticate by Workload Identity.
One identity with secret accessThe External Secrets Operator's. Application pods read synced k8s Secrets and have no Secret Manager grant at all.
Pod-to-pod least privilegeNetworkPolicies on every workload, enforced (Autopilot runs Dataplane V2).

Intra-cluster mTLS is NOT part of this pilot, and that is a decision. A service mesh (Istio / Cilium mTLS / Linkerd) would encrypt and mutually authenticate pod-to-pod traffic. It is not here because: the traffic never leaves the node network, which is itself inside a private VPC; a mesh adds a sidecar per pod (roughly +40% Autopilot cost on this footprint) and a second control plane to upgrade; and the threat it closes — an attacker who already has pod-network access — is one that NetworkPolicy already narrows to the specific ports each service needs. If media.net's security review requires it, GKE's managed Anthos Service Mesh is the path, and it is a scoped follow-on rather than a retrofit. Saying this out loud is the point: "TLS everywhere" means every hop that leaves a node, not every hop.

1.2 Cost model​

Method, so the numbers can be checked rather than believed. Autopilot bills the sum of pod requests, so the compute figure is derived from the resources.requests in deployments/overlays/mnet — 6.75 vCPU and 13.5 GiB across 16 pods — times 730 hours. (It was 6.0 / 12.0 / 13 before the architect review added the approval scheduler and the knowledge MCP server.) Managed-service prices are asia-south1 (Mumbai) list, retrieved from the GCP pricing pages 2026-09-05. Re-derive against the GCP pricing calculator before quoting this to the tenant; list prices move and this table cannot notice.

Line itemSizingEst. $/month
GKE Autopilot — cluster fee1 regional cluster74
GKE Autopilot — pod requests6.75 vCPU + 13.5 GiB, 730 h322
CloudSQL PG16 — REGIONAL HAdb-custom-2-8192, 2× for HA205
CloudSQL — storage + backups50 GB SSD (regional) + 30 backups25
Memorystore Redis, classic STANDARD_HA1 GB, primary + replica51
Cloud NAT1 gateway + ~200 GB processed41
GCS~100 GB standard + operations8
Vertex AI — embeddings~100 M tokens/mo, gemini-embedding-00115
Cloud Logging / Monitoringbeyond the free tier20
Secret Manager~20 secrets + access ops5
Load balancingnone — the tunnel replaces it0
Total infrastructure~766
Headroom to the $3000 budget~2234

THE MOST IMPORTANT LINE IN THIS TABLE IS THE ONE THAT IS NOT IN IT. LLM inference is BYOK: the Model Gateway calls Anthropic and OpenAI with the tenant's own provider keys, and that spend is billed by those vendors, not by GCP. So the $3000 GCP budget is not a cap on the product's cost — it is a cap on the infrastructure's cost, and at ~$729 it has 4× headroom.

Two consequences worth stating plainly:

  • The budget alerts become an anomaly detector rather than a spend ceiling. At this footprint, $1500 is more than double steady state; if it fires, something is wrong (a runaway job, a misconfigured retention, an unexpected egress) rather than merely busy. That is a better use of the alert than a threshold that fires every month and gets muted.
  • If the tenant routes LLM traffic through Vertex (Claude on Vertex, or Gemini) instead of direct-to-vendor, inference lands inside this bill and can dominate every line above it. Decide that before setting the budget, not after.

The actual cap on inference spend is in the product, not in GCP: the Model Gateway's per-org/team/agent spend gate. A GCP budget alerts, it does not stop.


2. Prerequisites (media.net supplies these)​

InputWhere it goes
GCP project idgcp:project, upsquad:hubProject, and every REPLACE_ME_MNET_PROJECT
Billing account id, with roles/billing.costsManager for the deployerupsquad:budget.billingAccount
Public hostname (e.g. upsquad.media.net)REPLACE_ME_MNET_APP_HOSTNAME
Cloudflare account id + an API tokenupsquad:cloudflare.accountId, CLOUDFLARE_API_TOKEN env at deploy time
Google Workspace IdP id from their Cloudflare Accessupsquad:cloudflare.allowedIdps
Clerk instance (publishable key, secret key, JWKS URL, two webhook secrets)Secret Manager (§5)
Confirmation that 10.30–10.34/16 do not overlap their corporate rangesupsquad:network

Confirm the CIDR question before the first apply. An overlap surfaces as intermittent unreachability from the WARP side, which is the worst failure signature to debug after the fact.

2.1 Bootstrap floor (a human, once, in their project)​

Pulumi cannot create the backend it stores its own state in.

PROJECT=<their-project>
gsutil mb -p "$PROJECT" -l asia-south1 "gs://${PROJECT}-pulumi-state"
gsutil versioning set on "gs://${PROJECT}-pulumi-state"
gcloud kms keyrings create pulumi --location asia-south1 --project "$PROJECT"
gcloud kms keys create state --location asia-south1 --keyring pulumi \
--purpose encryption --project "$PROJECT"
pulumi login "gs://${PROJECT}-pulumi-state"

Then edit Pulumi.mnet.yaml: replace secretsprovider with their KMS key. It ships pointed at a placeholder rather than at ours on purpose — a tenant's encrypted config must not be decryptable with our key.


3. Apply the infrastructure​

cd infra/pulumi/upsquad-infra
npm ci
npx tsc --noEmit # typecheck before touching a cloud
export CLOUDFLARE_API_TOKEN=... # never in the stack file
pulumi stack select mnet
pulumi preview --stack mnet # READ IT. Nothing below is reversible cheaply.
pulumi up --stack mnet

What is protected in Pulumi state (pulumi state unprotect required to remove): the VPC, the node subnet, the Artifact Registry repo (n/a here), the CloudSQL instance, the Memorystore cluster, the GCS bucket, and the GKE cluster. A pulumi destroy aimed at the wrong stack cannot take the tenant's data.

Record the outputs — the preflight and the overlay need them:

pulumi stack output --stack mnet platform # cluster name, endpoint, workload pool
gcloud sql instances describe cloudsql-mnet --project "$PROJECT" \
--format='value(settings.ipConfiguration.pscConfig)' # the PSC endpoint IP
gcloud redis clusters describe redis-mnet --region asia-south1 --project "$PROJECT" \
--format='value(pscConnections)' # the Redis endpoint

4. Cluster access, and the operator's first surprise​

4.1 The control plane has no public endpoint​

enablePrivateEndpoint: true with an empty masterAuthorizedCidrs means there is no external path to the API server at all. This is deliberate (§1.1) and it has a cost that is easier to accept before it is discovered:

  • Through the tunnel — add a Cloudflare Access application for the control plane and connect with cloudflared access tcp. This is the path that needs no extra GCP resource.
  • From a bastion — a small VM in the node subnet with gcloud container clusters get-credentials. Costs ~$7/month and is the option most SRE teams already have muscle memory for.

If media.net would rather allow their office egress range, add it to upsquad:gke.masterAuthorizedCidrs as a named entry. That is a deliberate widening with an owner, which is a different thing from a 0.0.0.0/0 nobody chose.

4.2 Install the External Secrets Operator​

It is a CRD-backed operator with its own upgrade cadence, so it is not in the overlay. Until it is installed, kubectl apply -k fails wholesale with no matches for kind ExternalSecret — loud, and correct.

kubectl create namespace platform
helm repo add external-secrets https://charts.external-secrets.io
helm install external-secrets external-secrets/external-secrets \
-n platform --set serviceAccount.name=external-secrets

The namespace must match the Workload Identity binding. upsquad:workloadIdentity.namespace in Pulumi.mnet.yaml is platform, so the binding names <project>.svc.id.goog[platform/external-secrets]. The Helm chart's default is its own external-secrets namespace; installing there without changing the Pulumi value gives a token exchange that 403s, and the symptom is every ExternalSecret stuck in SecretSyncedError.


5. Secret inventory​

Every credential lives in Secret Manager in the tenant's project. The ESO identity is the only one with secretmanager.secretAccessor.

Secret idWho writes itWhat reads it
database-url-mnetPulumi (DataStack)— (superseded by mnet-app-database-url)
redis-url-mnetPulumi (DataStack)CE, AO, MG — rediss://default:<auth>@host:port
redis-addr-mnetPulumi (DataStack)memory-mcp, knowledge-mcp
redis-password-mnetPulumi (DataStack)memory-mcp, knowledge-mcp
db-password-mnetPulumi (DataStack)the db Jobs (PGPASSWORD)
cloudflared-tunnel-token-mnetPulumi (CloudflareStack)the cloudflared Deployment
mnet-app-database-urloperatorCE, AO, MG, memory-mcp — the app_rw DSN
mnet-migrate-database-urloperatorthe migrate Job — the owner DSN
mnet-migration-db-roleoperatorevery service — names the role dbguard refuses
mnet-app-db-user / mnet-app-db-passwordoperatorthe db-roles Job
mnet-pghost / mnet-pguser / mnet-pgdatabaseoperatorthe db Jobs
mnet-vault-passphraseoperatorCE, AO, MG — unlocks the BYOK vault
mnet-layer-signing-keyoperatorCE
mnet-platform-signing-key-idoperator (printed by the key script)CE, AO
mnet-org-idoperatorCE (Clerk webhook default org)
mnet-clerk-jwks-urloperatorCE
mnet-clerk-publishable-key / mnet-clerk-secret-keyoperatorthe portal
mnet-clerk-gateway-webhook-secretoperatorCE
mnet-clerk-portal-webhook-secretoperatorthe portal
mnet-wikijs-api-keyoperator (from media.net)wikijs-mcp — see §7.2

mnet-wikijs-api-key is a WRITE credential even though nothing writes with it. Wiki.js 2.x has no read-only permission that covers get_page, so a key able to serve that tool necessarily holds manage:pages. The server issues no GraphQL mutation and internal/wikijsmcp/client_test.go asserts that at the wire — but size the blast radius from the KEY, not from the code. It is why the ingress NetworkPolicy in §7.2 is a release blocker.

The two Clerk webhook secrets are not interchangeable. Clerk issues one signing secret per endpoint, and this install registers two: the gateway's /webhooks/clerk and the portal's own Next.js route. Swapping them turns every delivery into a 401 that reads like a signature bug.

Create one:

gcloud secrets create mnet-vault-passphrase --project "$PROJECT" \
--replication-policy=user-managed --locations=asia-south1 \
--data-file=- <<< "$(openssl rand -base64 32)"

5.1 The platform signing key​

scripts/tenant/gen-platform-signing-key.sh --target k8s
# then publish the printed kid, exactly as the script instructs

Its absence is fail-safe and therefore silent: no key means the issuer routes are never mounted, governed MCP mints nothing, and nothing logs an error. The smoke test's JWKS probe is the only thing that tells you.


6. Image access​

Images are in our private GHCR. Every pod needs a pull secret:

kubectl -n platform create secret docker-registry ghcr-pull \
--docker-server=ghcr.io \
--docker-username=<github-user-or-bot> \
--docker-password=<GHCR read:packages token> \
--docker-email=devops@upsquad.ai

6.1 GHCR token vs. docker save / load​

GHCR deploy tokendocker save → transfer → docker load
Rotationone secret, rotate on a schedulenone to rotate
AvailabilityGHCR is in the tenant's deploy path — an outage blocks scaling and node replacementfully self-contained
Egressneeds ghcr.io on the allowlistnone
Effort per releasezeromanual, per image, per release
Provenancedigest-pinned by the registrywhatever landed on the USB stick

Use the token for this pilot. save/load is the fallback for an air-gapped install and is genuinely worse for a connected one: it turns every release into a manual operation, and manual operations are where "which build is running" stops being answerable.

The hardening follow-up is the third option and it is the right end state: mirror the images into an Artifact Registry inside the tenant's project and pull with Workload Identity. That removes the long-lived token entirely and takes GHCR out of the availability path. It needs a registry component whose IAM does not reach into our hub (RegistryStack's does), so it is scoped work rather than a config change.

The migrations image has no CI publish leg yet:

docker build -f deployments/db-jobs/Dockerfile \
-t ghcr.io/upsquad-ai/upsquad-migrations:main-<sha> . # from the repo root
docker push ghcr.io/upsquad-ai/upsquad-migrations:main-<sha>

7. Deploy and verify​

Fill every REPLACE_ME_* input in deployments/overlays/mnet. Derive the list rather than reading it from a sentence — this line used to enumerate six and the set has grown twice since:

grep -rho 'REPLACE_ME_[A-Z_]*' deployments/overlays/mnet | sort -u

As of core#3019 that is the project id, the app hostname, the org UUID, the two image tags, the node CIDR, the CloudSQL PSC IP, the Memorystore PSC CIDR, and media.net's Wiki.js base URL plus the CIDR its egress rule opens (§7.2). Then:

scripts/tenant/tenant-preflight.sh --mode k8s --namespace platform

It refuses on any unfilled placeholder, a non-false auth flag, a non-production ENVIRONMENT, a moving image tag, a missing signing key or pull secret, missing ESO CRDs, or an unconfirmed Cloud SQL RLS posture (§9.2).

kubectl apply -k deployments/overlays/mnet

# The three database Jobs are ORDER-DEPENDENT and Kubernetes will not enforce it
# for a manual apply — the sync-wave annotations only bind under ArgoCD.
kubectl -n platform wait --for=condition=complete job/upsquad-migrate --timeout=20m
kubectl -n platform wait --for=condition=complete job/upsquad-db-roles --timeout=5m
kubectl -n platform wait --for=condition=complete job/upsquad-platform-flags --timeout=5m

kubectl -n platform rollout status deploy/context-engine

Configure the Cloudflare tunnel's public hostname → service map (the tunnel is remotely managed, so routing lives in their dashboard, not in git):

PathService
/upsquad.*, /grpc.*, /rpc/*http://context-engine:8080
/llm/v1/* → rewrite /v1/*http://model-gateway:8091
/mcp/*, /.well-known/*, /oauth/token, /webhooks/clerkhttp://context-engine:8080
/ (catch-all)http://client:3000

/ws/* IS NO LONGER A ROUTE, AND AN EXISTING ONE MUST BE DELETED BY HAND. It pointed at http://context-engine:8081, the WS chat listener, which #3585 deleted along with the k8s Service ws port and the NetworkPolicy rules that admitted 8081 — so a route configured to it now resolves to a Service port that does not exist. This tunnel is remotely managed: git cannot remove a dashboard route. On any install configured before #3585, delete the /ws/* row in the Cloudflare dashboard as part of the upgrade; leaving it is a route that reads as deployed with nothing behind it. infra/edge/envoy.tenant.yaml expresses the replacement — /ws/ is SEALED with a 410 ws_chat_decommissioned direct response rather than removed, so an integrator gets a stated answer instead of the portal's catch-all HTML. Chat is Connect server-streaming to LifecycleService under /rpc/.

infra/edge/envoy.tenant.yaml is otherwise the same route set expressed as config. When the two shapes disagree about a path, diff against it — it is the version that is in git and under CI.

/llm/<anything-but-v1> must 404. cmd/model-gateway serves unauthenticated /metrics, /readyz and /healthz on the same mux as /v1/; the Envoy config seals that namespace with a direct_response, and the Cloudflare route table must not leave it falling through to the portal. The smoke test checks it.

Then:

scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>

7.1 Enabling knowledge-mcp (it ships stopped)​

deployments/overlays/mnet sets upsquad-knowledge-mcp to replicas: 0, with its PodDisruptionBudget at minAvailable: 0. The service is otherwise fully wired — base, ConfigMap, ExternalSecret, NetworkPolicy, publish leg, smoke-test entry.

Why stopped, and why the PDB moves with it. Un-Ready pods are still pods: a minAvailable: 1 budget over two ImagePullBackOff or CrashLoopBackOff replicas computes disruptionsAllowed: 0 and blocks eviction. On Autopilot, node auto-upgrade and auto-repair are evictions — so one un-deployable service would freeze node maintenance for the entire cluster. At zero replicas there are no pods to evict and no budget to violate.

Both preconditions must hold before you enable it:

  1. #2999 is merged and its image is published. upsquad-knowledge-mcp has a leg in publish-images.yml, source-existence-guarded so it stays inert until cmd/upsquad-knowledge-mcp/ is on main. Confirm the tag exists in GHCR.
  2. That binary reads REDIS_TLS. REDIS_TLS: "true" is already in the ConfigMap and is inert until then — the same one-file fix cmd/upsquad-memory-mcp/redis.go received in this PR.

Then move both numbers together — enabling replicas while the binary still crash-loops recreates the freeze this avoids:

# in deployments/overlays/mnet/kustomization.yaml, the two `patches:` entries
# Deployment upsquad-knowledge-mcp replicas 0 -> 2
# PDB upsquad-knowledge-mcp minAvailable 0 -> 1
kubectl apply -k deployments/overlays/mnet
kubectl -n platform rollout status deploy/upsquad-knowledge-mcp
scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>

The smoke test reports a zero-replica service as "WIRED BUT NOT ENABLED" rather than as a silent pass, so the switched-off state is visible in every run until you flip it.

It asserts the JWKS is served, /oauth/token is mounted and refusing, /llm/v1/ reaches the gateway, /llm/<other> is sealed, every Deployment is at its replica count, every Job completed, the connected DB role is NOSUPERUSER + NOBYPASSRLS, the replicas are on distinct nodes, and the portal answers 200 throughout a deliberate pod deletion.

7.2 Enabling wikijs-mcp — and the ingress policy that gates it​

deployments/overlays/mnet sets upsquad-wikijs-mcp to replicas: 0 with its PodDisruptionBudget at minAvailable: 0. The reason differs from knowledge-mcp's: the binary is fine, but WIKIJS_URL and WIKIJS_API_KEY are required — cmd/wikijs-mcp/config.go exits non-zero at boot without either — and media.net has not yet supplied the wiki's base URL. Two CrashLoopBackOff pods under a minAvailable: 1 budget compute disruptionsAllowed: 0 and block eviction, which on Autopilot freezes node auto-upgrade and auto-repair for the whole cluster. At zero replicas there are no pods and no budget to violate.

The release blocker: gateway-only INGRESS​

The overlay sets WIKIJS_MCP_AUTH_MODE: "gateway". In that posture the server answers every tools/call without a credential — correctly, because the team gateway has already authenticated, authorised and audited the caller and forwards no token for a type:none upstream registration.

The entire authorisation boundary is therefore the INGRESS NetworkPolicy. Without it, anything that can route to :9102 reads the whole wiki through an API key that necessarily holds manage:pages.

This is not a hardening step to schedule later. It has to hold before the first pod serves a call, and nothing else will tell you it does not:

with the policywithout it
Boot linehealthyhealthy
Liveness / readinessgreengreen
Every request serveda legitimate, correctly-answered tools/calla legitimate, correctly-answered tools/call
Logsnothing unusualnothing unusual

The binary cannot close this gap — a pod cannot read the NetworkPolicy that admits traffic to it, and a self-dial would prove nothing about what the cluster's policy controller enforces. So it states the precondition on a WARN line at every gateway-mode boot instead:

kubectl -n platform logs deploy/upsquad-wikijs-mcp \
| grep ingress_networkpolicy_gateway_only

That line appears on every healthy gateway-mode boot. It tells you the precondition applies — never whether it holds.

Verify the policy is ADMITTING, not merely APPLIED​

Two states look identical in kubectl get networkpolicy and neither enforces anything: a podSelector matching no pod (the policy governs an empty set, so the pods it was meant to guard are subject to no policy at all), and a cluster with no NetworkPolicy controller (the object is accepted by the API server and enforced by nobody). Check by measurement:

# 1. The policy's podSelector actually selects the pods. A NON-EMPTY pod list is
# the check; the policy object's existence is not.
kubectl -n platform get networkpolicy upsquad-wikijs-mcp -o yaml | grep -A4 podSelector
kubectl -n platform get pods -l app=upsquad-wikijs-mcp,component=mcp-server

# 2. There IS a policy controller. On GKE Autopilot network policy is always on
# (Dataplane V2); this is a real check only on a cluster somebody else built.
kubectl get nodes -o jsonpath='{.items[0].metadata.labels}' \
| tr ',' '\n' | grep -i dataplane

3. The measurement. From a pod that is not the gateway, the dial must HANG and then time out — a NetworkPolicy denial sends no RST and no ICMP. Use any existing pod in the namespace that has a shell; only if none does, run a throwaway image for the duration of the probe:

kubectl -n platform run np-probe --rm -it --restart=Never \
--image=curlimages/curl:8.10.1 -- \
curl -sS --max-time 8 http://upsquad-wikijs-mcp:9102/ ; echo "rc=$?"
  • rc=28 (timed out) — the boundary is enforcing. This is the pass.
  • rc=0, or any HTTP status body — STOP. The boundary is not there. Do not enable the service; anything that can route to the port is reading the wiki.
  • connection refused — nothing is listening. That is a different answer and it proves nothing about the policy; re-run once pods are Ready.

Both preconditions before you enable it​

  1. The image is published. wikijs-mcp has a leg in publish-images.yml producing ghcr.io/upsquad-ai/upsquad-wikijs-mcp. Confirm the pinned tag exists in GHCR — images: in the overlay is REPLACE_ME_IMAGE_TAG, never a moving tag.
  2. media.net has supplied the wiki inputs, and all three describe the same host:
    • WIKIJS_URL in configmaps.yaml — including the scheme;
    • REPLACE_ME_WIKIJS_EGRESS_CIDR in patches/networkpolicy-wikijs.yaml. A private-address wiki cannot be reached without this: the base's egress rule excepts RFC1918. PRD #2995 v1.2 records the same constraint for the ingestion arm (TK-3.6) and asks them for the base URL and whether it is publicly resolvable — the second half of that answer is this value;
    • mnet-wikijs-api-key in Secret Manager — a Wiki.js API key whose group holds read:pages + read:source, plus manage:pages or delete:pages for get_page.

Then move both numbers together:

# in deployments/overlays/mnet/kustomization.yaml, the two `patches:` entries
# Deployment upsquad-wikijs-mcp replicas 0 -> 2
# PDB upsquad-wikijs-mcp minAvailable 0 -> 1
kubectl apply -k deployments/overlays/mnet
kubectl -n platform rollout status deploy/upsquad-wikijs-mcp
# Re-run the ADMITTING measurement above, now that there are pods to select.
scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>

Register it on the team gateway (ADR-0024)​

A Ready Deployment does not make the wiki reachable to an agent — registration is a separate, data-plane step:

  1. Seed the egress row (ADR-0025) with allow_private and this Service's ClusterIP CIDR. Without it the gateway refuses to dial a private upstream.
  2. Register with transport streamable_http at http://upsquad-wikijs-mcp.platform.svc.cluster.local:9102/ and no upstream credential (type:none) — that is what makes gateway the correct WIKIJS_MCP_AUTH_MODE. The T1 probe calls initialize then tools/list; both stay open in either mode, so registration needs no token.
  3. Attach it to the units that should reach it. The gateway's clearance floor and CEL guardrails apply at the attachment; this server applies none of its own.
  4. Verify with a real tools/call. search_wiki answering no wiki page matches … on a populated wiki means the API key's group cannot see those pages — check the group's page rules, not this server.

Two failure shapes worth recognising​

SymptomCause
Pods Ready; every tool call returns "could not read the wiki" after ~20 sEgress. REPLACE_ME_WIKIJS_EGRESS_CIDR does not cover the wiki, or the wiki is http:// and only 443 is open. A NetworkPolicy denial sends no RST, so it presents as a timeout and nothing logs the network. Check this rule before you check the wiki.
Pods never become Ready; describe pod names the readiness probeIngress. The kubelet probes from the node address; REPLACE_ME_NODE_CIDR is wrong or unfilled. This is the base failing closed on purpose — unlike its siblings it ships no blanket ipBlock: 0.0.0.0/0 probe rule, because on this service that rule would be the hole rather than a convenience.

Rotating the API key​

gcloud secrets versions add mnet-wikijs-api-key --project "$PROJECT" --data-file=-
kubectl -n platform annotate externalsecret upsquad-wikijs-mcp-secrets \
force-sync=$(date +%s) --overwrite
kubectl -n platform rollout restart deploy/upsquad-wikijs-mcp

The restart is not optional. A Secret consumed through envFrom is snapshotted into the container's environment at start, so resyncing the Kubernetes Secret alone changes nothing the process can see — the old key keeps working until something else happens to restart the pod.


8. Rollback​

Every path is under five minutes (core tenet 5).

What brokeRollback
A bad imagekubectl -n platform set image deploy/<name> <container>=<previous digest> — or edit images: in the overlay and re-apply
A bad config valueEdit the ConfigMap, kubectl -n platform rollout restart deploy/<name>
A bad secretgcloud secrets versions add the previous value; ESO resyncs within refreshInterval (1 h) — force it with kubectl -n platform annotate externalsecret <n> force-sync=$(date +%s) --overwrite
app_rw breaks a code path (§9.1)Point mnet-app-database-url at the owner, set DB_ALLOW_RLS_BYPASS_ROLE=1, restart. Trades runtime tenant isolation for availability — deliberate, visible, and temporary.
Governed MCP misbehavingMCP_GATEWAY_ENABLED=false in the CE and AO ConfigMaps, restart both. They must move together.
The whole workload setkubectl delete -k deployments/overlays/mnet — leaves the data layer untouched
A bad migrationdocs/runbooks/dev-migrate-dirty-recovery.md. A schema rollback is data-destructive; it is a founder decision, not an operator one.

What is NOT a rollback: pulumi destroy. The data layer is protected in state and deletion-protected in the API, on purpose.

8.1 Failover drill (scheduled, not automatic)​

The smoke test deliberately does not trigger this — a regional-HA failover is disruptive (30–120 s) and belongs in a window. Run it once, before go-live. An HA pair nobody has ever failed over is a hypothesis.

# 1. Start a load generator against https://<hostname>/ in another terminal.
# 2. Fail over:
gcloud sql instances failover cloudsql-mnet --project "$PROJECT"
# 3. Record: time to first error, time to recovery, and whether any request
# returned a 5xx rather than a retry.

Expect a connection-error window. What you are measuring is whether the services reconnect rather than crash-loop, and how long the window is. Record the number on the install ticket; it is the answer to "what is our RPO/RTO for the database" and it is otherwise guesswork.


9. Known risks, named rather than discovered​

9.1 The RLS-bypass ledger (app_rw repoint)​

Services connect as app_rw (NOSUPERUSER NOBYPASSRLS) so FORCE RLS applies at runtime. 38 statements across 34 files still execute SET LOCAL row_security = off, and each hard-errors under that role. They are enumerated in test/lint/row_security_off_ledger.txt — read the file, not this sentence; it is derived and this number is not.

Most are in subsystems this spoke does not run (compliance retention, SIEM, SCIM, RTD, preview-router). Not all: internal/gateway/claims_enrich.go and internal/orgunit/* are on live paths. A staging boot is where this gets measured; the ledger is the map. The rollback is one secret version (§8).

9.2 Cloud SQL and the migration role (gate)​

docs/runbooks/migration-role-rls-posture.md records an unresolved question: Cloud SQL's cloudsqlsuperuser is documented without the BYPASSRLS attribute, and only a role that has it may grant it. If that is accurate, the #2781 migration-role mechanism is unavailable on Cloud SQL entirely, and a migration that touches tenant-scoped rows reports success having affected none of them — DELETE 0, exit 0, no error.

THE INSTALL CANNOT PROCEED ON A FALSE ASSUMPTION, which is what makes this a schedule risk rather than a correctness one: migration 214 RAISEs and stops the chain when the posture is wrong. The migrate Job runs as the owner and does not pretend otherwise.

RUN THE EXPERIMENT AS THE FIRST PRE-INSTALL TASK — before the install window opens, in media.net's project:

scripts/tenant/cloudsql-bypassrls-experiment.sh --project <TENANT_PROJECT>

It stands up a throwaway db-f1-micro, measures whether cloudsqlsuperuser can create or grant BYPASSRLS, prints the result, and deletes the instance — under an hour and under $1. It runs in their project rather than ours because our own GCP is founder-suspended for a deliberate spend cut.

Record the output on #2781 and on the install ticket, then set TENANT_RLS_POSTURE_CONFIRMED=1 for the preflight.

If the answer is "cannot", the failure branch is the 118-table alternative in migration-role-rls-posture.md. That is a founder call, not an operator one, and it should be raised the day the result is known — discovering it during a two-week Day-0 window is the expensive way to find out.

9.3 Egress, and the #2776 perimeter finding​

This install makes outbound calls:

DestinationPurposeAvoidable?
ghcr.ioimage pullsyes — mirror to Artifact Registry (§6.1)
ClerkJWKS + backend APIno — it is the identity provider
*.googleapis.comVertex, GCS, Secret Manager, loggingno (and it stays inside GCP)
Cloudflare edge (7844/443)the tunnelno — it is the ingress
Anthropic / OpenAIBYOK LLM calls via the Model Gatewayonly by not configuring a key
GitLab / YouTrack / Grafana / Gmail / Granolaregistered MCP serversper-connector
media.net's Wiki.jswikijs-mcp reads the SRE runbook corpusyes — by not enabling the service

The wiki egress is the one destination on this list that is NARROWED BY A CIDR RATHER THAN BOUNDED BY A CHOKEPOINT. Everything above rides the "internet on 443 except RFC1918" rule and is bounded by Cloud NAT and the VPC firewall; REPLACE_ME_WIKIJS_EGRESS_CIDR names one destination explicitly. Filling it with 0.0.0.0/0 is legal, is sometimes the only correct answer (a CDN-fronted wiki), and gives that rule the same shape as the others — record which answer was chosen on the install ticket. It is also the one egress whose destination is the tenant's own system rather than a third-party SaaS, so it is usually the easiest to pin to a /32.

#2776 — an on-prem install cannot be bounded to its perimeter by code. The gateway registers the hosted providers unconditionally, and the only thing preventing egress today is the absence of a BYOK key — a statement about a vault's contents, not a control. The actor who can write that key is a tenant/org administrator; the actor who owns the perimeter is the operator. On a tenant install that is a privilege inversion.

The host firewall is the control, and it is the tenant's to configure. On this spoke that means Cloud NAT plus the VPC firewall rules (infra/pulumi/upsquad-infra/firewall.ts) — every pod's egress passes through the NAT, so it is the one chokepoint their network team can log and restrict. Narrowing the Model Gateway's NetworkPolicy egress to a provider allowlist is the complementary in-cluster step and is a named follow-up.

#2776 was originally written against cmd/ai-gateway, which T18 (#2920) has since deleted outright — so the specific surface it named is gone everywhere, not merely undeployed on this spoke. The general property it describes (no code-level perimeter bound) still holds for the Model Gateway, which is the only LLM egress plane now, and a security questionnaire will ask about it. The host firewall remains the control.

9.4 Follow-ups this install ships without​

GapWhy it is not herePriority
knowledge-mcp ships WIRED AND STOPPED (replicas: 0)Two independent blockers: (1) cmd/upsquad-knowledge-mcp/ is not on main yet (#2999), so no image exists; (2) its binary (main.go:109) builds its Redis client with no TLSConfig against a TLS-only Memorystore, and there REDIS_ADDR is required and retrieval.NewService panics on nil — so it CrashLoopBackOffs rather than degrading. See §7.1 to enable.blocking, for that service only
wikijs-mcp ships WIRED AND STOPPED (replicas: 0)Not a code blocker — the binary is on main (#3017) and its image publishes. It is waiting on TENANT INPUT: media.net has not supplied the wiki's base URL, whether it is publicly resolvable, or an API key, and WIKIJS_URL/WIKIJS_API_KEY are both required at boot. Its gateway-only ingress NetworkPolicy is a release blocker for enabling it — see §7.2.blocking, for that service only
Redis IAM auth (#3006)This spoke uses the classic instance's static AUTH string because no binary can present an IAM token. IAM auth is the durable answer and spans four binaries.high
Artifact Registry mirror + Workload Identity pullNeeds a registry component whose IAM does not reach our hubhigh
Binary Authorization at admissionRequires an attestation pipeline; enabling it with images pulled from someone else's registry would make the cluster refuse its own workloadsmedium
PgBouncerServices connect straight to CloudSQL; connection counts are bounded by MAX_DB_CONNS × replicas rather than by a poolermedium
Custom-metrics adapterThe Model Gateway HPA's inflight metric has no source; it scales on CPU onlymedium
Intra-cluster mTLS§1.1low for the pilot
CI publish leg for upsquad-migrationspublish-images.yml is being edited by other work in flightlow
Tenant Prometheus alert rulesThe dev ruleset assumes node_exporter and the dev topologylow

10. The single-VM compose profile (secondary)​

docker-compose.tenant.yml + infra/edge/envoy.tenant.yaml + deploy/tenant/** + .env.tenant.example. Same hardening deltas as above, expressed for one machine: an in-container Envoy on :443/:80 (the only non-loopback publish), fake-gcs on a filesystem backend with a named volume, a bundled Ollama embedder (in-perimeter, keyless), and all the same pinned-shut auth flags.

cp .env.tenant.example /etc/upsquad/tenant.env && chmod 0600 /etc/upsquad/tenant.env
# fill every CHANGE_ME_
scripts/tenant/gen-platform-signing-key.sh --target compose
scripts/tenant/tenant-preflight.sh --mode compose --env-file /etc/upsquad/tenant.env
docker compose --env-file /etc/upsquad/tenant.env -f docker-compose.tenant.yml up -d
scripts/tenant/tenant-smoke.sh --mode compose --base-url https://<hostname>

Two differences from the GCP spoke worth knowing before quoting it:

  • Embeddings are local (bundled nomic-embed-text on Ollama, 768-d, keyless), not Vertex. That is a stronger perimeter claim — inference happens on hardware the operator controls — and a weaker performance one.
  • fake-gcs is the object store. Back up the gcsdata volume. It has no replication and no versioning.

The pgbouncer userlist is not in the repo (the dev one carries literal dev passwords under an explicit localhost-only exception that does not extend to a tenant's VM). Generate it:

mkdir -p /etc/upsquad/secrets/pgbouncer
printf '"app_rw" "%s"\n' "$APP_DB_PASSWORD" > /etc/upsquad/secrets/pgbouncer/userlist.txt
chmod 0600 /etc/upsquad/secrets/pgbouncer/userlist.txt

  • docs/runbooks/migration-role-rls-posture.md — §9.2's gate
  • docs/runbooks/embedding-dimension-change.md — before changing EMBEDDING_DIM
  • docs/runbooks/platform-token-issuer.md — the trust root
  • docs/runbooks/dev-migrate-dirty-recovery.md — a failed migration
  • infra/pulumi/upsquad-infra/README.md — the Pulumi program
  • scripts/check-model-gateway-exposure.py — what keeps :8091 off an unintended surface, in both install shapes
  • scripts/check-netpol-reachability.py — what keeps a NetworkPolicy peer from silently selecting zero pods. It renders deployments/upsquad-wikijs-mcp/base by name, because that base's ingress rule is §7.2's boundary rather than an ordinary flow
  • cmd/wikijs-mcp/README.md — the wikijs-mcp env contract, its auth-mode coherence gate, and the credential-hygiene argument behind §7.2