Runbook: Rehoming the Clerk webhook to the core gateway (F7)
Context: upsquad-client#288. The Clerk webhook (svix → MemberService sync)
used to live in the Next portal at src/app/api/webhooks/clerk/route.ts. A
static SPA cannot host server endpoints, so the webhook now lives on the core
gateway at POST /webhooks/clerk (internal/gateway/clerkwebhook).
This runbook covers the zero-gap cutover from the portal endpoint to the gateway endpoint. The founder performs the Clerk dashboard steps (agents cannot access the Clerk dashboard).
Enablement gates — all four must be crossed
The handler was merged, mounted, and unit-tested months before a single event could reach it, because four independent gates each masked the others: fixing any one changes nothing observable. Check them in order.
| Gate | What | Where | Status after core#2634 |
|---|---|---|---|
| G1 | handler exists | internal/gateway/clerkwebhook/ | ✅ merged (PR #1240) |
| G2 | mounted on the public (unauthenticated) mux | cmd/context-engine/main.go → RegisterPublicRoutes | ✅ merged |
| G3 | signing secret + default org id set | docker-compose.dev.yml ← repo-root .env | ⚙️ plumbed — founder supplies the values |
| G4 | Envoy edge route for /webhooks/clerk | infra/edge/envoy.yaml, dev/envoy/envoy.yaml | ✅ merged |
| G5 | Cloudflare Access bypass for the path | scripts/edge/cf-access-clerk-webhook-bypass.sh | 📝 script committed — founder runs it |
| G6 | Clerk dashboard dual-registration | Clerk dashboard | 📝 founder step — see below |
| G7 | portal endpoint removed | upsquad-client | ⛔ blocked on G6 |
Failure signatures — read the response, not the symptom
A webhook mis-registered against an SPA host does not look broken to the
sender: nginx's history-API fallback answers GET with 200 text/html. Here
it happens to be 405 for POST, which is the only reason this was caught by
inspection rather than by a silent month of dropped user.created events.
Never accept a 2xx alone as proof. Match the exact response to the gate:
| Response | Gate | Meaning |
|---|---|---|
302 → upsquad.cloudflareaccess.com/cdn-cgi/access/login/… | G5 | Cloudflare Access is gating the path. Clerk cannot log in; run the bypass script. |
405 text/html … nginx/… | G4 | No Envoy route — the SPA's static nginx answered. The request never reached the gateway. |
403 {"code":"Forbidden","message":"missing X-Tenant-Id header (ICS shim…)"} | G3 | Reached the gateway, but the handler is not mounted, so the auth catch-all took it. Looks like an auth bug; it is a missing env var. |
400 {"error":"missing svix headers"} | — | ✅ Handler is live and verifying. |
401 {"error":"invalid signature"} | — | ✅ Handler is live; wrong/absent signing secret for this endpoint. |
200 | — | ✅ Delivered. Still confirm the DB row (below). |
Reproduce any of these against the box with a simulated delivery:
# Through the public edge (exercises G5 + G4 + G3):
curl -sS -o /dev/stderr -w '\nHTTP %{http_code} %{content_type}\n' \
-X POST https://beta-dev.app.upsquad.ai/webhooks/clerk \
-H 'content-type: application/json' \
-H 'svix-id: msg_test' -H 'svix-timestamp: '"$(date +%s)" \
-H 'svix-signature: v1,dGVzdA==' \
--data '{"type":"user.created"}'
# Bypassing Cloudflare, straight at the host edge (isolates G4 + G3 from G5):
curl -sSk --http2 -o /dev/stderr -w '\nHTTP %{http_code} %{content_type}\n' \
-X POST https://127.0.0.1:10443/webhooks/clerk \
-H 'Host: beta-dev.app.upsquad.ai' -H 'content-type: application/json' \
-H 'svix-id: msg_test' -H 'svix-timestamp: '"$(date +%s)" \
-H 'svix-signature: v1,dGVzdA==' \
--data '{"type":"user.created"}'
A signature of 400/401 from the second command means everything inside the
box is correct and any remaining failure is Cloudflare Access (G5).
What the gateway endpoint does
- Verifies the svix signature over the raw body using
CLERK_WEBHOOK_SIGNING_SECRET(HMAC-SHA256,{id}.{timestamp}.{body}, 5-minute timestamp tolerance to defeat replay). No svix SDK dependency. - Routes events to MemberService inside an org-scoped DB transaction (RLS GUCs
set the same way the gateway Scope middleware does):
user.created→CreateMember(lookup-first, idempotent)user.updated→UpdateMember(upserts if the member is unknown)user.deleted→ deferred (logs + no-op; awaitsDisableUser, core#821)organizationMembership.created→AddTeamMembershiporganizationMembership.deleted→RemoveTeamMembership
- Acks
200on benign duplicates (AlreadyExists/NotFound); returns500on genuine failures so Clerk retries;401on signature failure;400on missing svix headers / malformed body.
Configuration (fail-closed)
The endpoint is OFF (not mounted) unless BOTH env vars are set on context-engine. A half-configured deploy logs loudly at startup and stays off so member-sync traffic is never silently dropped.
| Env var (in the container) | Purpose |
|---|---|
CLERK_WEBHOOK_SIGNING_SECRET | svix signing secret (whsec_...) from the Clerk dashboard endpoint config. Security gate — without it the webhook is not mounted. |
CLERK_WEBHOOK_DEFAULT_ORG_ID | UpsQuad org id that Clerk events route to (single-tenant fallback). Multi-tenant routing is a follow-up to core#1242. |
⚠️ The two receivers need two different secrets
Clerk issues a separate signing secret per endpoint. During dual
registration the portal and the gateway are both registered, so they hold
different whsec_ values at the same time. Giving the gateway the portal's
secret produces 401 invalid signature on every delivery — which reads like a
signature bug and sends you looking in the wrong place.
The host-side variable names are therefore deliberately distinct, and
docker-compose.dev.yml maps them per service:
Host variable (repo-root .env) | Container | Consumer |
|---|---|---|
CLERK_WEBHOOK_SIGNING_SECRET | client | portal route src/app/api/webhooks/clerk/route.ts |
CLERK_GATEWAY_WEBHOOK_SIGNING_SECRET | context-engine as CLERK_WEBHOOK_SIGNING_SECRET | gateway route POST /webhooks/clerk |
CLERK_WEBHOOK_DEFAULT_ORG_ID | context-engine | gateway event routing |
There is intentionally no fallback from the gateway variable to the portal one. Populate them with:
bash scripts/dev-secrets/drop-clerk.sh \
--publishable-key "pk_test_…" \
--secret-key "sk_test_…" \
--webhook-secret "whsec_…" `# portal endpoint` \
--gateway-webhook-secret "whsec_…" `# /webhooks/clerk endpoint` \
--default-org-id "00000000-0000-0000-0000-000000000001"
docker compose -f docker-compose.dev.yml up -d --force-recreate \
client context-engine
The script refuses a half-configured pair and refuses two identical secrets, so the fail-closed conditions are caught before the container restarts.
Verify at startup: look for Clerk webhook endpoint mounted path=/webhooks/clerk
in the context-engine logs. A DISABLED ERROR line names the missing var.
docker logs upsquad-context-engine 2>&1 | grep -i 'clerk webhook'
Dual-registration cutover (NO sync gap)
The gateway endpoint is designed to coexist with the portal endpoint so both receive deliveries during the transition. Sequence:
Step 1 — install the Envoy route (G4)
Merged in core#2634. Apply it to the running host edge:
scripts/edge/install-envoy-config.sh --dry-run # validate first
scripts/edge/install-envoy-config.sh # install + restart envoy
Confirm with the second curl in the failure-signature section above. You want
403 (route works, handler not yet configured) — not 405 nginx.
Step 2 — configure the receiver (G3)
Run drop-clerk.sh as shown in Configuration above, then recreate
context-engine and confirm the mounted log line. The signing secret comes
from step 4 below, so on a first pass you may only have the portal's secret —
in that case do step 4 first, copy the new endpoint's secret, then come back.
The 403 → 401 transition is how you know this step landed.
Step 3 — carve the Cloudflare Access bypass (G5)
beta-dev.app.upsquad.ai is Access-gated and Clerk cannot authenticate to
Cloudflare Access — no interactive identity, and no way to attach service-token
headers. Without this carve every delivery 302s forever.
scripts/edge/cf-access-clerk-webhook-bypass.sh # read-only check first
scripts/edge/cf-access-clerk-webhook-bypass.sh --apply # founder-run
The script is idempotent and self-verifying: it re-probes the public hostname
afterwards and maps whatever it gets back to the gate table above. To undo:
scripts/edge/cf-access-clerk-webhook-bypass.sh --revert.
Read the WHY-it-is-safe argument in the script header before running it — this is a deliberate public-edge auth reduction whose security then rests entirely on svix signature verification.
The carve is a path PREFIX on the Cloudflare side, not an exact path. Measured against the existing portal carve:
/api/webhooks/clerk→405and/api/webhooks/clerk/SUBPATH→404(both inside the carve), while/api/webhooks/OTHER→302to the Access login. Envoy's route, by contrast, matches exactly — so the un-gated subtree would have fallen through to the SPA catch-all and served the app shell publicly (GET /webhooks/clerk/anything→200 text/html, measured before the fix).Both Envoy configs therefore carry a paired
prefix: "/webhooks/clerk/"→direct_response: { status: 404 }rule immediately after the exact route. That rule is load-bearing for the Access bypass. If it is ever removed, revert the carve as well — the script's--verifyasserts the subtree returns404for exactly this reason, and the commentedbeta/app-prodpromotion stubs carry the pair so a promotion cannot reopen the hole.
Step 4 — dual-register in the Clerk dashboard (G6)
Clerk dashboard → Webhooks → Add Endpoint.
| Field | Value |
|---|---|
| Endpoint URL | https://beta-dev.app.upsquad.ai/webhooks/clerk |
| Subscribe to events | user.created, user.updated, user.deleted, organizationMembership.created, organizationMembership.deleted |
| Signing secret | auto-generated — reveal it and copy into CLERK_GATEWAY_WEBHOOK_SIGNING_SECRET (step 2) |
Exactly those five events: they are the ones events.go dispatches. Anything
else is accepted and ignored (200), so over-subscribing is harmless noise but
under-subscribing silently loses syncs.
Leave the existing portal endpoint registered. Both fire during the window; that is the point. The receivers are idempotent, so duplicate deliveries are safe — a repeat
user.createdno-ops when the member exists.
Step 5 — verify with a delivery AND a row (never a 2xx alone)
- Clerk dashboard → the new endpoint → Send test event, and/or trigger a
real sign-up. The endpoint's delivery log must show
200. - Cross-check the receiver actually did work:
docker logs upsquad-context-engine 2>&1 | grep -i 'clerk-webhook'
- Confirm the database row. This is the step that distinguishes "the
endpoint returned 200" from "member sync works":
docker exec upsquad-postgres psql -U upsquad -d upsquad -c \"SELECT id, email, member_type, created_at FROM membersWHERE org_id = '00000000-0000-0000-0000-000000000001'ORDER BY created_at DESC LIMIT 5;"
Step 6 — retire the portal registration (G7)
Only after step 5 passes, remove (or disable) the portal endpoint
registration in the Clerk dashboard. The portal route keeps existing until the
portal itself is retired; deregistering it in Clerk simply stops delivery.
Then upsquad-client can delete src/app/api/webhooks/clerk/route.ts
(upsquad-client#288 G7) and the /api/webhooks rule can come out of
dev/envoy/envoy.yaml.
Rollback
If the gateway endpoint misbehaves, re-enable the portal endpoint in the Clerk dashboard (it is unchanged) and disable the new one. No code rollback needed — both paths are idempotent against MemberService.
Full teardown, in reverse dependency order (each step is independently safe):
| Undo | Command / action |
|---|---|
| G6 | Clerk dashboard → delete the /webhooks/clerk endpoint; ensure the portal endpoint is enabled |
| G5 | scripts/edge/cf-access-clerk-webhook-bypass.sh --revert |
| G3 | remove CLERK_GATEWAY_WEBHOOK_SIGNING_SECRET + CLERK_WEBHOOK_DEFAULT_ORG_ID from .env, recreate context-engine (endpoint unmounts, logs the quiet INFO line) |
| G4 | revert the route in infra/edge/envoy.yaml, re-run scripts/edge/install-envoy-config.sh |
Reverting G5 alone is enough to stop external deliveries immediately; the rest are cleanup.
Known follow-ups
- Multi-tenant routing:
user.*events carry no UpsQuad tenant id, so today we route toCLERK_WEBHOOK_DEFAULT_ORG_ID. AGetOrgByClerkId-style resolution (or per-tenant Clerk endpoints) is tracked against core#1242. user.deletedpropagation waits onMemberService.DisableUser(core#821, OB-B3) for audit-preserving disable.