Skip to main content

Runbook: Rehoming the Clerk webhook to the core gateway (F7)

Context: upsquad-client#288. The Clerk webhook (svix → MemberService sync) used to live in the Next portal at src/app/api/webhooks/clerk/route.ts. A static SPA cannot host server endpoints, so the webhook now lives on the core gateway at POST /webhooks/clerk (internal/gateway/clerkwebhook).

This runbook covers the zero-gap cutover from the portal endpoint to the gateway endpoint. The founder performs the Clerk dashboard steps (agents cannot access the Clerk dashboard).

Enablement gates — all four must be crossed​

The handler was merged, mounted, and unit-tested months before a single event could reach it, because four independent gates each masked the others: fixing any one changes nothing observable. Check them in order.

GateWhatWhereStatus after core#2634
G1handler existsinternal/gateway/clerkwebhook/✅ merged (PR #1240)
G2mounted on the public (unauthenticated) muxcmd/context-engine/main.go → RegisterPublicRoutes✅ merged
G3signing secret + default org id setdocker-compose.dev.yml ← repo-root .env⚙️ plumbed — founder supplies the values
G4Envoy edge route for /webhooks/clerkinfra/edge/envoy.yaml, dev/envoy/envoy.yaml✅ merged
G5Cloudflare Access bypass for the pathscripts/edge/cf-access-clerk-webhook-bypass.sh📝 script committed — founder runs it
G6Clerk dashboard dual-registrationClerk dashboard📝 founder step — see below
G7portal endpoint removedupsquad-client⛔ blocked on G6

Failure signatures — read the response, not the symptom​

A webhook mis-registered against an SPA host does not look broken to the sender: nginx's history-API fallback answers GET with 200 text/html. Here it happens to be 405 for POST, which is the only reason this was caught by inspection rather than by a silent month of dropped user.created events.

Never accept a 2xx alone as proof. Match the exact response to the gate:

ResponseGateMeaning
302 → upsquad.cloudflareaccess.com/cdn-cgi/access/login/…G5Cloudflare Access is gating the path. Clerk cannot log in; run the bypass script.
405 text/html … nginx/…G4No Envoy route — the SPA's static nginx answered. The request never reached the gateway.
403 {"code":"Forbidden","message":"missing X-Tenant-Id header (ICS shim…)"}G3Reached the gateway, but the handler is not mounted, so the auth catch-all took it. Looks like an auth bug; it is a missing env var.
400 {"error":"missing svix headers"}—✅ Handler is live and verifying.
401 {"error":"invalid signature"}—✅ Handler is live; wrong/absent signing secret for this endpoint.
200—✅ Delivered. Still confirm the DB row (below).

Reproduce any of these against the box with a simulated delivery:

# Through the public edge (exercises G5 + G4 + G3):
curl -sS -o /dev/stderr -w '\nHTTP %{http_code} %{content_type}\n' \
-X POST https://beta-dev.app.upsquad.ai/webhooks/clerk \
-H 'content-type: application/json' \
-H 'svix-id: msg_test' -H 'svix-timestamp: '"$(date +%s)" \
-H 'svix-signature: v1,dGVzdA==' \
--data '{"type":"user.created"}'

# Bypassing Cloudflare, straight at the host edge (isolates G4 + G3 from G5):
curl -sSk --http2 -o /dev/stderr -w '\nHTTP %{http_code} %{content_type}\n' \
-X POST https://127.0.0.1:10443/webhooks/clerk \
-H 'Host: beta-dev.app.upsquad.ai' -H 'content-type: application/json' \
-H 'svix-id: msg_test' -H 'svix-timestamp: '"$(date +%s)" \
-H 'svix-signature: v1,dGVzdA==' \
--data '{"type":"user.created"}'

A signature of 400/401 from the second command means everything inside the box is correct and any remaining failure is Cloudflare Access (G5).

What the gateway endpoint does​

  • Verifies the svix signature over the raw body using CLERK_WEBHOOK_SIGNING_SECRET (HMAC-SHA256, {id}.{timestamp}.{body}, 5-minute timestamp tolerance to defeat replay). No svix SDK dependency.
  • Routes events to MemberService inside an org-scoped DB transaction (RLS GUCs set the same way the gateway Scope middleware does):
    • user.created → CreateMember (lookup-first, idempotent)
    • user.updated → UpdateMember (upserts if the member is unknown)
    • user.deleted → deferred (logs + no-op; awaits DisableUser, core#821)
    • organizationMembership.created → AddTeamMembership
    • organizationMembership.deleted → RemoveTeamMembership
  • Acks 200 on benign duplicates (AlreadyExists / NotFound); returns 500 on genuine failures so Clerk retries; 401 on signature failure; 400 on missing svix headers / malformed body.

Configuration (fail-closed)​

The endpoint is OFF (not mounted) unless BOTH env vars are set on context-engine. A half-configured deploy logs loudly at startup and stays off so member-sync traffic is never silently dropped.

Env var (in the container)Purpose
CLERK_WEBHOOK_SIGNING_SECRETsvix signing secret (whsec_...) from the Clerk dashboard endpoint config. Security gate — without it the webhook is not mounted.
CLERK_WEBHOOK_DEFAULT_ORG_IDUpsQuad org id that Clerk events route to (single-tenant fallback). Multi-tenant routing is a follow-up to core#1242.

⚠️ The two receivers need two different secrets​

Clerk issues a separate signing secret per endpoint. During dual registration the portal and the gateway are both registered, so they hold different whsec_ values at the same time. Giving the gateway the portal's secret produces 401 invalid signature on every delivery — which reads like a signature bug and sends you looking in the wrong place.

The host-side variable names are therefore deliberately distinct, and docker-compose.dev.yml maps them per service:

Host variable (repo-root .env)ContainerConsumer
CLERK_WEBHOOK_SIGNING_SECRETclientportal route src/app/api/webhooks/clerk/route.ts
CLERK_GATEWAY_WEBHOOK_SIGNING_SECRETcontext-engine as CLERK_WEBHOOK_SIGNING_SECRETgateway route POST /webhooks/clerk
CLERK_WEBHOOK_DEFAULT_ORG_IDcontext-enginegateway event routing

There is intentionally no fallback from the gateway variable to the portal one. Populate them with:

bash scripts/dev-secrets/drop-clerk.sh \
--publishable-key "pk_test_…" \
--secret-key "sk_test_…" \
--webhook-secret "whsec_…" `# portal endpoint` \
--gateway-webhook-secret "whsec_…" `# /webhooks/clerk endpoint` \
--default-org-id "00000000-0000-0000-0000-000000000001"

docker compose -f docker-compose.dev.yml up -d --force-recreate \
client context-engine

The script refuses a half-configured pair and refuses two identical secrets, so the fail-closed conditions are caught before the container restarts.

Verify at startup: look for Clerk webhook endpoint mounted path=/webhooks/clerk in the context-engine logs. A DISABLED ERROR line names the missing var.

docker logs upsquad-context-engine 2>&1 | grep -i 'clerk webhook'

Dual-registration cutover (NO sync gap)​

The gateway endpoint is designed to coexist with the portal endpoint so both receive deliveries during the transition. Sequence:

Step 1 — install the Envoy route (G4)​

Merged in core#2634. Apply it to the running host edge:

scripts/edge/install-envoy-config.sh --dry-run # validate first
scripts/edge/install-envoy-config.sh # install + restart envoy

Confirm with the second curl in the failure-signature section above. You want 403 (route works, handler not yet configured) — not 405 nginx.

Step 2 — configure the receiver (G3)​

Run drop-clerk.sh as shown in Configuration above, then recreate context-engine and confirm the mounted log line. The signing secret comes from step 4 below, so on a first pass you may only have the portal's secret — in that case do step 4 first, copy the new endpoint's secret, then come back. The 403 → 401 transition is how you know this step landed.

Step 3 — carve the Cloudflare Access bypass (G5)​

beta-dev.app.upsquad.ai is Access-gated and Clerk cannot authenticate to Cloudflare Access — no interactive identity, and no way to attach service-token headers. Without this carve every delivery 302s forever.

scripts/edge/cf-access-clerk-webhook-bypass.sh # read-only check first
scripts/edge/cf-access-clerk-webhook-bypass.sh --apply # founder-run

The script is idempotent and self-verifying: it re-probes the public hostname afterwards and maps whatever it gets back to the gate table above. To undo: scripts/edge/cf-access-clerk-webhook-bypass.sh --revert.

Read the WHY-it-is-safe argument in the script header before running it — this is a deliberate public-edge auth reduction whose security then rests entirely on svix signature verification.

The carve is a path PREFIX on the Cloudflare side, not an exact path. Measured against the existing portal carve: /api/webhooks/clerk → 405 and /api/webhooks/clerk/SUBPATH → 404 (both inside the carve), while /api/webhooks/OTHER → 302 to the Access login. Envoy's route, by contrast, matches exactly — so the un-gated subtree would have fallen through to the SPA catch-all and served the app shell publicly (GET /webhooks/clerk/anything → 200 text/html, measured before the fix).

Both Envoy configs therefore carry a paired prefix: "/webhooks/clerk/" → direct_response: { status: 404 } rule immediately after the exact route. That rule is load-bearing for the Access bypass. If it is ever removed, revert the carve as well — the script's --verify asserts the subtree returns 404 for exactly this reason, and the commented beta / app-prod promotion stubs carry the pair so a promotion cannot reopen the hole.

Step 4 — dual-register in the Clerk dashboard (G6)​

Clerk dashboard → Webhooks → Add Endpoint.

FieldValue
Endpoint URLhttps://beta-dev.app.upsquad.ai/webhooks/clerk
Subscribe to eventsuser.created, user.updated, user.deleted, organizationMembership.created, organizationMembership.deleted
Signing secretauto-generated — reveal it and copy into CLERK_GATEWAY_WEBHOOK_SIGNING_SECRET (step 2)

Exactly those five events: they are the ones events.go dispatches. Anything else is accepted and ignored (200), so over-subscribing is harmless noise but under-subscribing silently loses syncs.

Leave the existing portal endpoint registered. Both fire during the window; that is the point. The receivers are idempotent, so duplicate deliveries are safe — a repeat user.created no-ops when the member exists.

Step 5 — verify with a delivery AND a row (never a 2xx alone)​

  1. Clerk dashboard → the new endpoint → Send test event, and/or trigger a real sign-up. The endpoint's delivery log must show 200.
  2. Cross-check the receiver actually did work:
    docker logs upsquad-context-engine 2>&1 | grep -i 'clerk-webhook'
  3. Confirm the database row. This is the step that distinguishes "the endpoint returned 200" from "member sync works":
    docker exec upsquad-postgres psql -U upsquad -d upsquad -c \
    "SELECT id, email, member_type, created_at FROM members
    WHERE org_id = '00000000-0000-0000-0000-000000000001'
    ORDER BY created_at DESC LIMIT 5;"

Step 6 — retire the portal registration (G7)​

Only after step 5 passes, remove (or disable) the portal endpoint registration in the Clerk dashboard. The portal route keeps existing until the portal itself is retired; deregistering it in Clerk simply stops delivery. Then upsquad-client can delete src/app/api/webhooks/clerk/route.ts (upsquad-client#288 G7) and the /api/webhooks rule can come out of dev/envoy/envoy.yaml.

Rollback​

If the gateway endpoint misbehaves, re-enable the portal endpoint in the Clerk dashboard (it is unchanged) and disable the new one. No code rollback needed — both paths are idempotent against MemberService.

Full teardown, in reverse dependency order (each step is independently safe):

UndoCommand / action
G6Clerk dashboard → delete the /webhooks/clerk endpoint; ensure the portal endpoint is enabled
G5scripts/edge/cf-access-clerk-webhook-bypass.sh --revert
G3remove CLERK_GATEWAY_WEBHOOK_SIGNING_SECRET + CLERK_WEBHOOK_DEFAULT_ORG_ID from .env, recreate context-engine (endpoint unmounts, logs the quiet INFO line)
G4revert the route in infra/edge/envoy.yaml, re-run scripts/edge/install-envoy-config.sh

Reverting G5 alone is enough to stop external deliveries immediately; the rest are cleanup.

Known follow-ups​

  • Multi-tenant routing: user.* events carry no UpsQuad tenant id, so today we route to CLERK_WEBHOOK_DEFAULT_ORG_ID. A GetOrgByClerkId-style resolution (or per-tenant Clerk endpoints) is tracked against core#1242.
  • user.deleted propagation waits on MemberService.DisableUser (core#821, OB-B3) for audit-preserving disable.