Tenant install — the media.net GCP spoke
Issue core#2992, tracker #2991. ADR-0018 (cloud foundation), ADR-0017 (staged promotion), ADR-0034 (edge posture).
Read this first. This runbook describes two install shapes. They are not alternatives of equal standing:
Shape Status Where GCP spoke — private regional GKE Autopilot in the tenant's own project PRIMARY. The media.net deliverable. §1–§9 Single-VM compose — one hardened docker composeprofileSECONDARY. Generic on-prem artefact, kept for future prospects. §10 The compose profile was built first and was largely complete when the founder pivoted the target to a GCP project (2026-09-05). It is finished and shipped rather than deleted because a self-contained single-VM install has standalone value — but nothing about media.net's delivery runs on it.
1. What gets built, and what it costs to run
Two artefacts, applied in order:
-
infra/pulumi/upsquad-infra, stackmnet— the tenant's project: private VPC, Cloud NAT, regional private GKE Autopilot cluster, CloudSQL PostgreSQL 16 with regional HA and private-IP-only + enforced TLS, Memorystore Redis with a replica and in-transit encryption, a GCS bucket, Secret Manager, Workload Identity, a Cloudflare tunnel, and a budget. -
deployments/overlays/mnet— the workloads: Context Engine, Model Gateway, Agent Orchestrator, the memory MCP server, the knowledge MCP server, the approval scheduler, the client portal, the cloudflared connector, and the three database Jobs.Two of those were added by the architect's review of PR #3004. The approval scheduler because without it overdue governance approvals never expire — #2556 measured 71 pending approvals, every one past its
expires_at, on a stack running without it. The knowledge MCP server because its base was always this PR's follow-through on #2993; note it is blocked from deploying until that binary gains TLS support for Redis (§9.4).
1.1 Security posture, stated as properties
| Property | How it is held |
|---|---|
| No inbound internet surface | No GCLB, no Service: LoadBalancer, no node external IPs, no Ingress object. Traffic arrives because the cloudflared pod dials out. |
| TLS at every off-node hop | Cloudflare edge TLS → tunnel (TLS) → in-cluster; CloudSQL sslMode: ENCRYPTED_ONLY; Memorystore SERVER_AUTHENTICATION + a static AUTH string; GCS native TLS. |
| Private control plane | enablePrivateEndpoint: true, masterAuthorizedCidrs: []. kubectl needs a path into the VPC (§4.1). |
| No auth bypass reachable | ENVIRONMENT: production makes the gateway's bypass triple-gate structurally unreachable; DISABLE_AUTH/ICS_MODE/NEXT_PUBLIC_DISABLE_AUTH are false and asserted by preflight. |
| Tenant isolation live at runtime | Services connect as app_rw (NOSUPERUSER NOBYPASSRLS), so FORCE RLS applies. Smoke-tested (§7). |
| No service-account key files | Vertex, GCS and the Drive connector authenticate by Workload Identity. |
| One identity with secret access | The External Secrets Operator's. Application pods read synced k8s Secrets and have no Secret Manager grant at all. |
| Pod-to-pod least privilege | NetworkPolicies on every workload, enforced (Autopilot runs Dataplane V2). |
Intra-cluster mTLS is NOT part of this pilot, and that is a decision. A service mesh (Istio / Cilium mTLS / Linkerd) would encrypt and mutually authenticate pod-to-pod traffic. It is not here because: the traffic never leaves the node network, which is itself inside a private VPC; a mesh adds a sidecar per pod (roughly +40% Autopilot cost on this footprint) and a second control plane to upgrade; and the threat it closes — an attacker who already has pod-network access — is one that NetworkPolicy already narrows to the specific ports each service needs. If media.net's security review requires it, GKE's managed Anthos Service Mesh is the path, and it is a scoped follow-on rather than a retrofit. Saying this out loud is the point: "TLS everywhere" means every hop that leaves a node, not every hop.
1.2 Cost model
Method, so the numbers can be checked rather than believed. Autopilot bills
the sum of pod requests, so the compute figure is derived from the
resources.requests in deployments/overlays/mnet — 6.75 vCPU and 13.5 GiB
across 16 pods — times 730 hours. (It was 6.0 / 12.0 / 13 before the architect
review added the approval scheduler and the knowledge MCP server.) Managed-service prices are asia-south1 (Mumbai)
list, retrieved from the GCP pricing pages 2026-09-05. Re-derive against the
GCP pricing calculator before quoting this to the tenant; list prices move and
this table cannot notice.
| Line item | Sizing | Est. $/month |
|---|---|---|
| GKE Autopilot — cluster fee | 1 regional cluster | 74 |
| GKE Autopilot — pod requests | 6.75 vCPU + 13.5 GiB, 730 h | 322 |
| CloudSQL PG16 — REGIONAL HA | db-custom-2-8192, 2× for HA | 205 |
| CloudSQL — storage + backups | 50 GB SSD (regional) + 30 backups | 25 |
Memorystore Redis, classic STANDARD_HA | 1 GB, primary + replica | 51 |
| Cloud NAT | 1 gateway + ~200 GB processed | 41 |
| GCS | ~100 GB standard + operations | 8 |
| Vertex AI — embeddings | ~100 M tokens/mo, gemini-embedding-001 | 15 |
| Cloud Logging / Monitoring | beyond the free tier | 20 |
| Secret Manager | ~20 secrets + access ops | 5 |
| Load balancing | none — the tunnel replaces it | 0 |
| Total infrastructure | ~766 | |
| Headroom to the $3000 budget | ~2234 |
THE MOST IMPORTANT LINE IN THIS TABLE IS THE ONE THAT IS NOT IN IT. LLM inference is BYOK: the Model Gateway calls Anthropic and OpenAI with the tenant's own provider keys, and that spend is billed by those vendors, not by GCP. So the $3000 GCP budget is not a cap on the product's cost — it is a cap on the infrastructure's cost, and at ~$729 it has 4× headroom.
Two consequences worth stating plainly:
- The budget alerts become an anomaly detector rather than a spend ceiling. At this footprint, $1500 is more than double steady state; if it fires, something is wrong (a runaway job, a misconfigured retention, an unexpected egress) rather than merely busy. That is a better use of the alert than a threshold that fires every month and gets muted.
- If the tenant routes LLM traffic through Vertex (Claude on Vertex, or Gemini) instead of direct-to-vendor, inference lands inside this bill and can dominate every line above it. Decide that before setting the budget, not after.
The actual cap on inference spend is in the product, not in GCP: the Model Gateway's per-org/team/agent spend gate. A GCP budget alerts, it does not stop.
2. Prerequisites (media.net supplies these)
| Input | Where it goes |
|---|---|
| GCP project id | gcp:project, upsquad:hubProject, and every REPLACE_ME_MNET_PROJECT |
Billing account id, with roles/billing.costsManager for the deployer | upsquad:budget.billingAccount |
Public hostname (e.g. upsquad.media.net) | REPLACE_ME_MNET_APP_HOSTNAME |
| Cloudflare account id + an API token | upsquad:cloudflare.accountId, CLOUDFLARE_API_TOKEN env at deploy time |
| Google Workspace IdP id from their Cloudflare Access | upsquad:cloudflare.allowedIdps |
| Clerk instance (publishable key, secret key, JWKS URL, two webhook secrets) | Secret Manager (§5) |
Confirmation that 10.30–10.34/16 do not overlap their corporate ranges | upsquad:network |
Confirm the CIDR question before the first apply. An overlap surfaces as intermittent unreachability from the WARP side, which is the worst failure signature to debug after the fact.
2.1 Bootstrap floor (a human, once, in their project)
Pulumi cannot create the backend it stores its own state in.
PROJECT=<their-project>
gsutil mb -p "$PROJECT" -l asia-south1 "gs://${PROJECT}-pulumi-state"
gsutil versioning set on "gs://${PROJECT}-pulumi-state"
gcloud kms keyrings create pulumi --location asia-south1 --project "$PROJECT"
gcloud kms keys create state --location asia-south1 --keyring pulumi \
--purpose encryption --project "$PROJECT"
pulumi login "gs://${PROJECT}-pulumi-state"
Then edit Pulumi.mnet.yaml: replace secretsprovider with their KMS key.
It ships pointed at a placeholder rather than at ours on purpose — a tenant's
encrypted config must not be decryptable with our key.
3. Apply the infrastructure
cd infra/pulumi/upsquad-infra
npm ci
npx tsc --noEmit # typecheck before touching a cloud
export CLOUDFLARE_API_TOKEN=... # never in the stack file
pulumi stack select mnet
pulumi preview --stack mnet # READ IT. Nothing below is reversible cheaply.
pulumi up --stack mnet
What is protected in Pulumi state (pulumi state unprotect required to
remove): the VPC, the node subnet, the Artifact Registry repo (n/a here), the
CloudSQL instance, the Memorystore cluster, the GCS bucket, and the GKE cluster.
A pulumi destroy aimed at the wrong stack cannot take the tenant's data.
Record the outputs — the preflight and the overlay need them:
pulumi stack output --stack mnet platform # cluster name, endpoint, workload pool
gcloud sql instances describe cloudsql-mnet --project "$PROJECT" \
--format='value(settings.ipConfiguration.pscConfig)' # the PSC endpoint IP
gcloud redis clusters describe redis-mnet --region asia-south1 --project "$PROJECT" \
--format='value(pscConnections)' # the Redis endpoint
4. Cluster access, and the operator's first surprise
4.1 The control plane has no public endpoint
enablePrivateEndpoint: true with an empty masterAuthorizedCidrs means
there is no external path to the API server at all. This is deliberate
(§1.1) and it has a cost that is easier to accept before it is discovered:
- Through the tunnel — add a Cloudflare Access application for the control
plane and connect with
cloudflared access tcp. This is the path that needs no extra GCP resource. - From a bastion — a small VM in the node subnet with
gcloud container clusters get-credentials. Costs ~$7/month and is the option most SRE teams already have muscle memory for.
If media.net would rather allow their office egress range, add it to
upsquad:gke.masterAuthorizedCidrs as a named entry. That is a deliberate
widening with an owner, which is a different thing from a 0.0.0.0/0 nobody
chose.
4.2 Install the External Secrets Operator
It is a CRD-backed operator with its own upgrade cadence, so it is not in
the overlay. Until it is installed, kubectl apply -k fails wholesale with
no matches for kind ExternalSecret — loud, and correct.
kubectl create namespace platform
helm repo add external-secrets https://charts.external-secrets.io
helm install external-secrets external-secrets/external-secrets \
-n platform --set serviceAccount.name=external-secrets
The namespace must match the Workload Identity binding.
upsquad:workloadIdentity.namespaceinPulumi.mnet.yamlisplatform, so the binding names<project>.svc.id.goog[platform/external-secrets]. The Helm chart's default is its ownexternal-secretsnamespace; installing there without changing the Pulumi value gives a token exchange that 403s, and the symptom is every ExternalSecret stuck inSecretSyncedError.
5. Secret inventory
Every credential lives in Secret Manager in the tenant's project. The ESO
identity is the only one with secretmanager.secretAccessor.
| Secret id | Who writes it | What reads it |
|---|---|---|
database-url-mnet | Pulumi (DataStack) | — (superseded by mnet-app-database-url) |
redis-url-mnet | Pulumi (DataStack) | CE, AO, MG — rediss://default:<auth>@host:port |
redis-addr-mnet | Pulumi (DataStack) | memory-mcp, knowledge-mcp |
redis-password-mnet | Pulumi (DataStack) | memory-mcp, knowledge-mcp |
db-password-mnet | Pulumi (DataStack) | the db Jobs (PGPASSWORD) |
cloudflared-tunnel-token-mnet | Pulumi (CloudflareStack) | the cloudflared Deployment |
mnet-app-database-url | operator | CE, AO, MG, memory-mcp — the app_rw DSN |
mnet-migrate-database-url | operator | the migrate Job — the owner DSN |
mnet-migration-db-role | operator | every service — names the role dbguard refuses |
mnet-app-db-user / mnet-app-db-password | operator | the db-roles Job |
mnet-pghost / mnet-pguser / mnet-pgdatabase | operator | the db Jobs |
mnet-vault-passphrase | operator | CE, AO, MG — unlocks the BYOK vault |
mnet-layer-signing-key | operator | CE |
mnet-platform-signing-key-id | operator (printed by the key script) | CE, AO |
mnet-org-id | operator | CE (Clerk webhook default org) |
mnet-clerk-jwks-url | operator | CE |
mnet-clerk-publishable-key / mnet-clerk-secret-key | operator | the portal |
mnet-clerk-gateway-webhook-secret | operator | CE |
mnet-clerk-portal-webhook-secret | operator | the portal |
mnet-wikijs-api-key | operator (from media.net) | wikijs-mcp — see §7.2 |
mnet-wikijs-api-keyis a WRITE credential even though nothing writes with it. Wiki.js 2.x has no read-only permission that coversget_page, so a key able to serve that tool necessarily holdsmanage:pages. The server issues no GraphQL mutation andinternal/wikijsmcp/client_test.goasserts that at the wire — but size the blast radius from the KEY, not from the code. It is why the ingress NetworkPolicy in §7.2 is a release blocker.
The two Clerk webhook secrets are not interchangeable. Clerk issues one signing secret per endpoint, and this install registers two: the gateway's
/webhooks/clerkand the portal's own Next.js route. Swapping them turns every delivery into a 401 that reads like a signature bug.
Create one:
gcloud secrets create mnet-vault-passphrase --project "$PROJECT" \
--replication-policy=user-managed --locations=asia-south1 \
--data-file=- <<< "$(openssl rand -base64 32)"
5.1 The platform signing key
scripts/tenant/gen-platform-signing-key.sh --target k8s
# then publish the printed kid, exactly as the script instructs
Its absence is fail-safe and therefore silent: no key means the issuer routes are never mounted, governed MCP mints nothing, and nothing logs an error. The smoke test's JWKS probe is the only thing that tells you.
6. Image access
Images are in our private GHCR. Every pod needs a pull secret:
kubectl -n platform create secret docker-registry ghcr-pull \
--docker-server=ghcr.io \
--docker-username=<github-user-or-bot> \
--docker-password=<GHCR read:packages token> \
--docker-email=devops@upsquad.ai
6.1 GHCR token vs. docker save / load
| GHCR deploy token | docker save → transfer → docker load | |
|---|---|---|
| Rotation | one secret, rotate on a schedule | none to rotate |
| Availability | GHCR is in the tenant's deploy path — an outage blocks scaling and node replacement | fully self-contained |
| Egress | needs ghcr.io on the allowlist | none |
| Effort per release | zero | manual, per image, per release |
| Provenance | digest-pinned by the registry | whatever landed on the USB stick |
Use the token for this pilot. save/load is the fallback for an air-gapped
install and is genuinely worse for a connected one: it turns every release into
a manual operation, and manual operations are where "which build is running"
stops being answerable.
The hardening follow-up is the third option and it is the right end state:
mirror the images into an Artifact Registry inside the tenant's project and
pull with Workload Identity. That removes the long-lived token entirely and
takes GHCR out of the availability path. It needs a registry component whose IAM
does not reach into our hub (RegistryStack's does), so it is scoped work
rather than a config change.
The migrations image has no CI publish leg yet:
docker build -f deployments/db-jobs/Dockerfile \
-t ghcr.io/upsquad-ai/upsquad-migrations:main-<sha> . # from the repo root
docker push ghcr.io/upsquad-ai/upsquad-migrations:main-<sha>
7. Deploy and verify
Fill every REPLACE_ME_* input in deployments/overlays/mnet. Derive the
list rather than reading it from a sentence — this line used to enumerate six
and the set has grown twice since:
grep -rho 'REPLACE_ME_[A-Z_]*' deployments/overlays/mnet | sort -u
As of core#3019 that is the project id, the app hostname, the org UUID, the two image tags, the node CIDR, the CloudSQL PSC IP, the Memorystore PSC CIDR, and media.net's Wiki.js base URL plus the CIDR its egress rule opens (§7.2). Then:
scripts/tenant/tenant-preflight.sh --mode k8s --namespace platform
It refuses on any unfilled placeholder, a non-false auth flag, a
non-production ENVIRONMENT, a moving image tag, a missing signing key or
pull secret, missing ESO CRDs, or an unconfirmed Cloud SQL RLS posture (§9.2).
kubectl apply -k deployments/overlays/mnet
# The three database Jobs are ORDER-DEPENDENT and Kubernetes will not enforce it
# for a manual apply — the sync-wave annotations only bind under ArgoCD.
kubectl -n platform wait --for=condition=complete job/upsquad-migrate --timeout=20m
kubectl -n platform wait --for=condition=complete job/upsquad-db-roles --timeout=5m
kubectl -n platform wait --for=condition=complete job/upsquad-platform-flags --timeout=5m
kubectl -n platform rollout status deploy/context-engine
Configure the Cloudflare tunnel's public hostname → service map (the tunnel is remotely managed, so routing lives in their dashboard, not in git):
| Path | Service |
|---|---|
/upsquad.*, /grpc.*, /rpc/* | http://context-engine:8080 |
/llm/v1/* → rewrite /v1/* | http://model-gateway:8091 |
/mcp/*, /.well-known/*, /oauth/token, /webhooks/clerk | http://context-engine:8080 |
/ (catch-all) | http://client:3000 |
/ws/*IS NO LONGER A ROUTE, AND AN EXISTING ONE MUST BE DELETED BY HAND. It pointed athttp://context-engine:8081, the WS chat listener, which #3585 deleted along with the k8s Servicewsport and the NetworkPolicy rules that admitted 8081 — so a route configured to it now resolves to a Service port that does not exist. This tunnel is remotely managed: git cannot remove a dashboard route. On any install configured before #3585, delete the/ws/*row in the Cloudflare dashboard as part of the upgrade; leaving it is a route that reads as deployed with nothing behind it.infra/edge/envoy.tenant.yamlexpresses the replacement —/ws/is SEALED with a410 ws_chat_decommissioneddirect response rather than removed, so an integrator gets a stated answer instead of the portal's catch-all HTML. Chat is Connect server-streaming toLifecycleServiceunder/rpc/.
infra/edge/envoy.tenant.yamlis otherwise the same route set expressed as config. When the two shapes disagree about a path, diff against it — it is the version that is in git and under CI.
/llm/<anything-but-v1>must 404.cmd/model-gatewayserves unauthenticated/metrics,/readyzand/healthzon the same mux as/v1/; the Envoy config seals that namespace with adirect_response, and the Cloudflare route table must not leave it falling through to the portal. The smoke test checks it.
Then:
scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>
7.1 Enabling knowledge-mcp (it ships stopped)
deployments/overlays/mnet sets upsquad-knowledge-mcp to replicas: 0,
with its PodDisruptionBudget at minAvailable: 0. The service is otherwise
fully wired — base, ConfigMap, ExternalSecret, NetworkPolicy, publish leg,
smoke-test entry.
Why stopped, and why the PDB moves with it. Un-Ready pods are still pods: a
minAvailable: 1 budget over two ImagePullBackOff or CrashLoopBackOff
replicas computes disruptionsAllowed: 0 and blocks eviction. On Autopilot,
node auto-upgrade and auto-repair are evictions — so one un-deployable service
would freeze node maintenance for the entire cluster. At zero replicas there are
no pods to evict and no budget to violate.
Both preconditions must hold before you enable it:
- #2999 is merged and its image is published.
upsquad-knowledge-mcphas a leg inpublish-images.yml, source-existence-guarded so it stays inert untilcmd/upsquad-knowledge-mcp/is on main. Confirm the tag exists in GHCR. - That binary reads
REDIS_TLS.REDIS_TLS: "true"is already in the ConfigMap and is inert until then — the same one-file fixcmd/upsquad-memory-mcp/redis.goreceived in this PR.
Then move both numbers together — enabling replicas while the binary still crash-loops recreates the freeze this avoids:
# in deployments/overlays/mnet/kustomization.yaml, the two `patches:` entries
# Deployment upsquad-knowledge-mcp replicas 0 -> 2
# PDB upsquad-knowledge-mcp minAvailable 0 -> 1
kubectl apply -k deployments/overlays/mnet
kubectl -n platform rollout status deploy/upsquad-knowledge-mcp
scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>
The smoke test reports a zero-replica service as "WIRED BUT NOT ENABLED" rather than as a silent pass, so the switched-off state is visible in every run until you flip it.
It asserts the JWKS is served, /oauth/token is mounted and refusing,
/llm/v1/ reaches the gateway, /llm/<other> is sealed, every Deployment is at
its replica count, every Job completed, the connected DB role is NOSUPERUSER +
NOBYPASSRLS, the replicas are on distinct nodes, and the portal answers 200
throughout a deliberate pod deletion.
7.2 Enabling wikijs-mcp — and the ingress policy that gates it
deployments/overlays/mnet sets upsquad-wikijs-mcp to replicas: 0 with
its PodDisruptionBudget at minAvailable: 0. The reason differs from
knowledge-mcp's: the binary is fine, but WIKIJS_URL and WIKIJS_API_KEY are
required — cmd/wikijs-mcp/config.go exits non-zero at boot without either —
and media.net has not yet supplied the wiki's base URL. Two CrashLoopBackOff
pods under a minAvailable: 1 budget compute disruptionsAllowed: 0 and
block eviction, which on Autopilot freezes node auto-upgrade and auto-repair
for the whole cluster. At zero replicas there are no pods and no budget to
violate.
The release blocker: gateway-only INGRESS
The overlay sets WIKIJS_MCP_AUTH_MODE: "gateway". In that posture the server
answers every tools/call without a credential — correctly, because the team
gateway has already authenticated, authorised and audited the caller and
forwards no token for a type:none upstream registration.
The entire authorisation boundary is therefore the INGRESS NetworkPolicy. Without it, anything that can route to
:9102reads the whole wiki through an API key that necessarily holdsmanage:pages.
This is not a hardening step to schedule later. It has to hold before the first pod serves a call, and nothing else will tell you it does not:
| with the policy | without it | |
|---|---|---|
| Boot line | healthy | healthy |
| Liveness / readiness | green | green |
| Every request served | a legitimate, correctly-answered tools/call | a legitimate, correctly-answered tools/call |
| Logs | nothing unusual | nothing unusual |
The binary cannot close this gap — a pod cannot read the NetworkPolicy that admits traffic to it, and a self-dial would prove nothing about what the cluster's policy controller enforces. So it states the precondition on a WARN line at every gateway-mode boot instead:
kubectl -n platform logs deploy/upsquad-wikijs-mcp \
| grep ingress_networkpolicy_gateway_only
That line appears on every healthy gateway-mode boot. It tells you the precondition applies — never whether it holds.
Verify the policy is ADMITTING, not merely APPLIED
Two states look identical in kubectl get networkpolicy and neither enforces
anything: a podSelector matching no pod (the policy governs an empty set, so
the pods it was meant to guard are subject to no policy at all), and a cluster
with no NetworkPolicy controller (the object is accepted by the API server and
enforced by nobody). Check by measurement:
# 1. The policy's podSelector actually selects the pods. A NON-EMPTY pod list is
# the check; the policy object's existence is not.
kubectl -n platform get networkpolicy upsquad-wikijs-mcp -o yaml | grep -A4 podSelector
kubectl -n platform get pods -l app=upsquad-wikijs-mcp,component=mcp-server
# 2. There IS a policy controller. On GKE Autopilot network policy is always on
# (Dataplane V2); this is a real check only on a cluster somebody else built.
kubectl get nodes -o jsonpath='{.items[0].metadata.labels}' \
| tr ',' '\n' | grep -i dataplane
3. The measurement. From a pod that is not the gateway, the dial must HANG and then time out — a NetworkPolicy denial sends no RST and no ICMP. Use any existing pod in the namespace that has a shell; only if none does, run a throwaway image for the duration of the probe:
kubectl -n platform run np-probe --rm -it --restart=Never \
--image=curlimages/curl:8.10.1 -- \
curl -sS --max-time 8 http://upsquad-wikijs-mcp:9102/ ; echo "rc=$?"
rc=28(timed out) — the boundary is enforcing. This is the pass.rc=0, or any HTTP status body — STOP. The boundary is not there. Do not enable the service; anything that can route to the port is reading the wiki.connection refused— nothing is listening. That is a different answer and it proves nothing about the policy; re-run once pods are Ready.
Both preconditions before you enable it
- The image is published.
wikijs-mcphas a leg inpublish-images.ymlproducingghcr.io/upsquad-ai/upsquad-wikijs-mcp. Confirm the pinned tag exists in GHCR —images:in the overlay isREPLACE_ME_IMAGE_TAG, never a moving tag. - media.net has supplied the wiki inputs, and all three describe the same
host:
WIKIJS_URLinconfigmaps.yaml— including the scheme;REPLACE_ME_WIKIJS_EGRESS_CIDRinpatches/networkpolicy-wikijs.yaml. A private-address wiki cannot be reached without this: the base's egress rule excepts RFC1918. PRD #2995 v1.2 records the same constraint for the ingestion arm (TK-3.6) and asks them for the base URL and whether it is publicly resolvable — the second half of that answer is this value;mnet-wikijs-api-keyin Secret Manager — a Wiki.js API key whose group holdsread:pages+read:source, plusmanage:pagesordelete:pagesforget_page.
Then move both numbers together:
# in deployments/overlays/mnet/kustomization.yaml, the two `patches:` entries
# Deployment upsquad-wikijs-mcp replicas 0 -> 2
# PDB upsquad-wikijs-mcp minAvailable 0 -> 1
kubectl apply -k deployments/overlays/mnet
kubectl -n platform rollout status deploy/upsquad-wikijs-mcp
# Re-run the ADMITTING measurement above, now that there are pods to select.
scripts/tenant/tenant-smoke.sh --mode k8s --base-url https://<hostname>
Register it on the team gateway (ADR-0024)
A Ready Deployment does not make the wiki reachable to an agent — registration is a separate, data-plane step:
- Seed the egress row (ADR-0025) with
allow_privateand this Service's ClusterIP CIDR. Without it the gateway refuses to dial a private upstream. - Register with transport
streamable_httpathttp://upsquad-wikijs-mcp.platform.svc.cluster.local:9102/and no upstream credential (type:none) — that is what makesgatewaythe correctWIKIJS_MCP_AUTH_MODE. The T1 probe callsinitializethentools/list; both stay open in either mode, so registration needs no token. - Attach it to the units that should reach it. The gateway's clearance floor and CEL guardrails apply at the attachment; this server applies none of its own.
- Verify with a real
tools/call.search_wikiansweringno wiki page matches …on a populated wiki means the API key's group cannot see those pages — check the group's page rules, not this server.
Two failure shapes worth recognising
| Symptom | Cause |
|---|---|
| Pods Ready; every tool call returns "could not read the wiki" after ~20 s | Egress. REPLACE_ME_WIKIJS_EGRESS_CIDR does not cover the wiki, or the wiki is http:// and only 443 is open. A NetworkPolicy denial sends no RST, so it presents as a timeout and nothing logs the network. Check this rule before you check the wiki. |
Pods never become Ready; describe pod names the readiness probe | Ingress. The kubelet probes from the node address; REPLACE_ME_NODE_CIDR is wrong or unfilled. This is the base failing closed on purpose — unlike its siblings it ships no blanket ipBlock: 0.0.0.0/0 probe rule, because on this service that rule would be the hole rather than a convenience. |
Rotating the API key
gcloud secrets versions add mnet-wikijs-api-key --project "$PROJECT" --data-file=-
kubectl -n platform annotate externalsecret upsquad-wikijs-mcp-secrets \
force-sync=$(date +%s) --overwrite
kubectl -n platform rollout restart deploy/upsquad-wikijs-mcp
The restart is not optional. A Secret consumed through envFrom is
snapshotted into the container's environment at start, so resyncing the
Kubernetes Secret alone changes nothing the process can see — the old key keeps
working until something else happens to restart the pod.
8. Rollback
Every path is under five minutes (core tenet 5).
| What broke | Rollback |
|---|---|
| A bad image | kubectl -n platform set image deploy/<name> <container>=<previous digest> — or edit images: in the overlay and re-apply |
| A bad config value | Edit the ConfigMap, kubectl -n platform rollout restart deploy/<name> |
| A bad secret | gcloud secrets versions add the previous value; ESO resyncs within refreshInterval (1 h) — force it with kubectl -n platform annotate externalsecret <n> force-sync=$(date +%s) --overwrite |
app_rw breaks a code path (§9.1) | Point mnet-app-database-url at the owner, set DB_ALLOW_RLS_BYPASS_ROLE=1, restart. Trades runtime tenant isolation for availability — deliberate, visible, and temporary. |
| Governed MCP misbehaving | MCP_GATEWAY_ENABLED=false in the CE and AO ConfigMaps, restart both. They must move together. |
| The whole workload set | kubectl delete -k deployments/overlays/mnet — leaves the data layer untouched |
| A bad migration | docs/runbooks/dev-migrate-dirty-recovery.md. A schema rollback is data-destructive; it is a founder decision, not an operator one. |
What is NOT a rollback: pulumi destroy. The data layer is protected in
state and deletion-protected in the API, on purpose.
8.1 Failover drill (scheduled, not automatic)
The smoke test deliberately does not trigger this — a regional-HA failover is disruptive (30–120 s) and belongs in a window. Run it once, before go-live. An HA pair nobody has ever failed over is a hypothesis.
# 1. Start a load generator against https://<hostname>/ in another terminal.
# 2. Fail over:
gcloud sql instances failover cloudsql-mnet --project "$PROJECT"
# 3. Record: time to first error, time to recovery, and whether any request
# returned a 5xx rather than a retry.
Expect a connection-error window. What you are measuring is whether the services reconnect rather than crash-loop, and how long the window is. Record the number on the install ticket; it is the answer to "what is our RPO/RTO for the database" and it is otherwise guesswork.
9. Known risks, named rather than discovered
9.1 The RLS-bypass ledger (app_rw repoint)
Services connect as app_rw (NOSUPERUSER NOBYPASSRLS) so FORCE RLS applies at
runtime. 38 statements across 34 files still execute
SET LOCAL row_security = off, and each hard-errors under that role. They are
enumerated in test/lint/row_security_off_ledger.txt — read the file, not this
sentence; it is derived and this number is not.
Most are in subsystems this spoke does not run (compliance retention, SIEM,
SCIM, RTD, preview-router). Not all: internal/gateway/claims_enrich.go and
internal/orgunit/* are on live paths. A staging boot is where this gets
measured; the ledger is the map. The rollback is one secret version (§8).
9.2 Cloud SQL and the migration role (gate)
docs/runbooks/migration-role-rls-posture.md records an unresolved question:
Cloud SQL's cloudsqlsuperuser is documented without the BYPASSRLS
attribute, and only a role that has it may grant it. If that is accurate, the
#2781 migration-role mechanism is unavailable on Cloud SQL entirely, and a
migration that touches tenant-scoped rows reports success having affected none
of them — DELETE 0, exit 0, no error.
THE INSTALL CANNOT PROCEED ON A FALSE ASSUMPTION, which is what makes this a schedule risk rather than a correctness one: migration 214 RAISEs and stops the chain when the posture is wrong. The migrate Job runs as the owner and does not pretend otherwise.
RUN THE EXPERIMENT AS THE FIRST PRE-INSTALL TASK — before the install window opens, in media.net's project:
scripts/tenant/cloudsql-bypassrls-experiment.sh --project <TENANT_PROJECT>
It stands up a throwaway db-f1-micro, measures whether cloudsqlsuperuser can
create or grant BYPASSRLS, prints the result, and deletes the instance —
under an hour and under $1. It runs in their project rather than ours
because our own GCP is founder-suspended for a deliberate spend cut.
Record the output on #2781 and on the install ticket, then set
TENANT_RLS_POSTURE_CONFIRMED=1 for the preflight.
If the answer is "cannot", the failure branch is the 118-table alternative
in migration-role-rls-posture.md. That is a founder call, not an operator
one, and it should be raised the day the result is known — discovering it during
a two-week Day-0 window is the expensive way to find out.
9.3 Egress, and the #2776 perimeter finding
This install makes outbound calls:
| Destination | Purpose | Avoidable? |
|---|---|---|
ghcr.io | image pulls | yes — mirror to Artifact Registry (§6.1) |
| Clerk | JWKS + backend API | no — it is the identity provider |
*.googleapis.com | Vertex, GCS, Secret Manager, logging | no (and it stays inside GCP) |
| Cloudflare edge (7844/443) | the tunnel | no — it is the ingress |
| Anthropic / OpenAI | BYOK LLM calls via the Model Gateway | only by not configuring a key |
| GitLab / YouTrack / Grafana / Gmail / Granola | registered MCP servers | per-connector |
| media.net's Wiki.js | wikijs-mcp reads the SRE runbook corpus | yes — by not enabling the service |
The wiki egress is the one destination on this list that is NARROWED BY A CIDR RATHER THAN BOUNDED BY A CHOKEPOINT. Everything above rides the "internet on 443 except RFC1918" rule and is bounded by Cloud NAT and the VPC firewall;
REPLACE_ME_WIKIJS_EGRESS_CIDRnames one destination explicitly. Filling it with0.0.0.0/0is legal, is sometimes the only correct answer (a CDN-fronted wiki), and gives that rule the same shape as the others — record which answer was chosen on the install ticket. It is also the one egress whose destination is the tenant's own system rather than a third-party SaaS, so it is usually the easiest to pin to a /32.
#2776 — an on-prem install cannot be bounded to its perimeter by code. The gateway registers the hosted providers unconditionally, and the only thing preventing egress today is the absence of a BYOK key — a statement about a vault's contents, not a control. The actor who can write that key is a tenant/org administrator; the actor who owns the perimeter is the operator. On a tenant install that is a privilege inversion.
The host firewall is the control, and it is the tenant's to configure. On this spoke that means Cloud NAT plus the VPC firewall rules (
infra/pulumi/upsquad-infra/firewall.ts) — every pod's egress passes through the NAT, so it is the one chokepoint their network team can log and restrict. Narrowing the Model Gateway's NetworkPolicy egress to a provider allowlist is the complementary in-cluster step and is a named follow-up.#2776 was originally written against
cmd/ai-gateway, which T18 (#2920) has since deleted outright — so the specific surface it named is gone everywhere, not merely undeployed on this spoke. The general property it describes (no code-level perimeter bound) still holds for the Model Gateway, which is the only LLM egress plane now, and a security questionnaire will ask about it. The host firewall remains the control.
9.4 Follow-ups this install ships without
| Gap | Why it is not here | Priority |
|---|---|---|
knowledge-mcp ships WIRED AND STOPPED (replicas: 0) | Two independent blockers: (1) cmd/upsquad-knowledge-mcp/ is not on main yet (#2999), so no image exists; (2) its binary (main.go:109) builds its Redis client with no TLSConfig against a TLS-only Memorystore, and there REDIS_ADDR is required and retrieval.NewService panics on nil — so it CrashLoopBackOffs rather than degrading. See §7.1 to enable. | blocking, for that service only |
wikijs-mcp ships WIRED AND STOPPED (replicas: 0) | Not a code blocker — the binary is on main (#3017) and its image publishes. It is waiting on TENANT INPUT: media.net has not supplied the wiki's base URL, whether it is publicly resolvable, or an API key, and WIKIJS_URL/WIKIJS_API_KEY are both required at boot. Its gateway-only ingress NetworkPolicy is a release blocker for enabling it — see §7.2. | blocking, for that service only |
| Redis IAM auth (#3006) | This spoke uses the classic instance's static AUTH string because no binary can present an IAM token. IAM auth is the durable answer and spans four binaries. | high |
| Artifact Registry mirror + Workload Identity pull | Needs a registry component whose IAM does not reach our hub | high |
| Binary Authorization at admission | Requires an attestation pipeline; enabling it with images pulled from someone else's registry would make the cluster refuse its own workloads | medium |
| PgBouncer | Services connect straight to CloudSQL; connection counts are bounded by MAX_DB_CONNS × replicas rather than by a pooler | medium |
| Custom-metrics adapter | The Model Gateway HPA's inflight metric has no source; it scales on CPU only | medium |
| Intra-cluster mTLS | §1.1 | low for the pilot |
CI publish leg for upsquad-migrations | publish-images.yml is being edited by other work in flight | low |
| Tenant Prometheus alert rules | The dev ruleset assumes node_exporter and the dev topology | low |
10. The single-VM compose profile (secondary)
docker-compose.tenant.yml + infra/edge/envoy.tenant.yaml +
deploy/tenant/** + .env.tenant.example. Same hardening deltas as above,
expressed for one machine: an in-container Envoy on :443/:80 (the only
non-loopback publish), fake-gcs on a filesystem backend with a named
volume, a bundled Ollama embedder (in-perimeter, keyless), and all the same
pinned-shut auth flags.
cp .env.tenant.example /etc/upsquad/tenant.env && chmod 0600 /etc/upsquad/tenant.env
# fill every CHANGE_ME_
scripts/tenant/gen-platform-signing-key.sh --target compose
scripts/tenant/tenant-preflight.sh --mode compose --env-file /etc/upsquad/tenant.env
docker compose --env-file /etc/upsquad/tenant.env -f docker-compose.tenant.yml up -d
scripts/tenant/tenant-smoke.sh --mode compose --base-url https://<hostname>
Two differences from the GCP spoke worth knowing before quoting it:
- Embeddings are local (bundled
nomic-embed-texton Ollama, 768-d, keyless), not Vertex. That is a stronger perimeter claim — inference happens on hardware the operator controls — and a weaker performance one. fake-gcsis the object store. Back up thegcsdatavolume. It has no replication and no versioning.
The pgbouncer userlist is not in the repo (the dev one carries literal dev passwords under an explicit localhost-only exception that does not extend to a tenant's VM). Generate it:
mkdir -p /etc/upsquad/secrets/pgbouncer
printf '"app_rw" "%s"\n' "$APP_DB_PASSWORD" > /etc/upsquad/secrets/pgbouncer/userlist.txt
chmod 0600 /etc/upsquad/secrets/pgbouncer/userlist.txt
11. Related
docs/runbooks/migration-role-rls-posture.md— §9.2's gatedocs/runbooks/embedding-dimension-change.md— before changingEMBEDDING_DIMdocs/runbooks/platform-token-issuer.md— the trust rootdocs/runbooks/dev-migrate-dirty-recovery.md— a failed migrationinfra/pulumi/upsquad-infra/README.md— the Pulumi programscripts/check-model-gateway-exposure.py— what keeps:8091off an unintended surface, in both install shapesscripts/check-netpol-reachability.py— what keeps a NetworkPolicy peer from silently selecting zero pods. It rendersdeployments/upsquad-wikijs-mcp/baseby name, because that base's ingress rule is §7.2's boundary rather than an ordinary flowcmd/wikijs-mcp/README.md— the wikijs-mcp env contract, its auth-mode coherence gate, and the credential-hygiene argument behind §7.2