KT — gwpoc: the Claude Code Inspection Gateway
Audience: an engineer (hi Ashik) testing or extending the gwpoc spike. Goal: understand the whole thing end-to-end — what it proves, how the pieces wire together, how to drive the two-lane demo, and exactly where the honest edges are.
Status legend used throughout: 🟢 LIVE (merged, running on the devbox) · 🟡 QUIRK/CAVEAT (works, with a sharp edge you should know) · 🔵 DEV-ONLY (depends on the beta-dev stack's trust model) · 🟠 FOUNDER-GATED/PARKED (exists but requires a human grant or decision).
This is a spike, not a product (
README.mdsays so in its header). It informs core#2616 (Claude Agent SDK as coder runtime), which isawaiting-approval— nothing here is an accepted platform direction. Treat every claim as "demonstrated on the devbox," not "shipped."
Canonical code: upsquad-client → tools/gateway-poc/ @ 49ca8e4e (PR client#804, 33 files, Go 1.25, zero external deps). The devbox operating copy /opt/upsquad/ops/gateway-poc carries the same 33 source files byte-for-byte (verified 2026-08-21; the devbox tree additionally holds untracked runtime state — captures, lanes, binaries — that never leaves the box). All file:line references below resolve in either tree.
1. What it is, in one paragraph
gwpoc serve is an inspecting pass-through proxy between an unmodified Claude Code and api.anthropic.com. Point the CLI at it with one env var (ANTHROPIC_BASE_URL=http://localhost:8787) — keys untouched — and every exchange is forwarded faithfully (SSE unbuffered, ~0.8 ms p50 overhead) while a structured record lands on disk: who (four identity carriers), what (masked full-prompt attestation + tool calls), and what it cost (all four token classes, cache-aware — a naive input-token meter misreads real sessions by ~10×). On top of that sit the two-lane demo: the same CLI run "vanilla" vs "upsquad" (a PreToolUse hook enforcing the live GovernanceService cascade + an MCP shim serving Context-Engine retrieval), with per-session verdicts (delivered / cut_off / …) computed mechanically from the wire format, side by side in a Compare UI.
2. At a glance
| Piece | What | Status |
|---|---|---|
gwpoc serve (serve.go) | the proxy + capture pipeline, :8787 loopback via systemd gwpoc.service | 🟢 v0.6.0, 3d+ uptime, NRestarts=0 |
Identity (identity.go) | 4 carriers, precedence header > path > keymap > metadata | 🟢 (keymap dormant — identities.json is comment-only by design) |
Capture (capture.go, mask.go) | append-only JSONL ledger (fsync/row) + masked body sidecars | 🟢 |
Pricing (pricing.go, reprice.go) | 20-row rate card + 2-row Fast-mode card; null-vs-true-$0 discipline; history reprice | 🟢 card snapshot "verified 2026-08-17" 🟡 |
Sessions & verdicts (sessions.go) | per-session rollup + outcome + governance denial counts | 🟢 |
/ui viewer (ui.go, ui.html) | live board, Compare tab, CSV export, typed-confirm Clear | 🟢 (no auth of its own — see §10) 🟡 |
demo/gwdemo.sh | the agent bootstrap script: lane assembly, shim lifecycle, run/live, exit rollups | 🟢 |
demo/hooks/govern-check.sh | PreToolUse → live GovernanceService/Check, fail-closed | 🟢 enforcement is client-side 🟡 |
gwpoc mcpshim (mcpshim.go) | upsquad_context MCP server: search_knowledge, get_context | 🟢 tool surface · 🔵 auth model |
Grant scripts (verify/grant-demo-*.sh) | seed the org-unit allows (reads + the two MCP tools) | 🟠 founder-run |
| Exposure | https://gw-dev.upsquad.ai/ui via CF tunnel + CF Access (@upsquad.ai) | 🟢 kill-switch: teardown-gw-exposure.sh |
| SDK-runtime decision | core#2616 | 🟠 awaiting-approval |
3. Architecture
Walkthrough (numbers = request path, letters = the upsquad lane's extra wiring):
0 — the viewer is public only through the CF tunnel; unauthenticated hits get a 302 to CF Access. Origin binds loopback-only.
1 — both lanes are stock Claude Code with ANTHROPIC_BASE_URL pointed at the gateway; identity/profile/session-name ride ANTHROPIC_CUSTOM_HEADERS.
2 — the proxy forwards method/body/query/headers verbatim minus RFC-7230 hop-by-hop and the four internal X-Upsquad-* headers (they never leave the building — serve.go:251-254); the one mutation is Accept-Encoding: gzip (serve.go:260) because an unparseable Brotli body blinds capture (FINDINGS #18). SSE flushes per event; delivery is never gated on capture.
3 — each exchange appends one JSONL row (fsync'd) + masked full-body sidecars. Raw credentials never persist — only sha256(cred)[:12].
4 — /ui re-scans the ledger on a 3 s poll; /ui and /favicon.ico are the only paths not forwarded (FINDINGS #15: observing must not perturb).
A/B — in the upsquad lane, a PreToolUse hook maps every tool to a governance target (tool:claude_code.Bash, tool:mcp.upsquad_context.search_knowledge) and asks the live platform cascade. Any error ⇒ deny (fail-closed).
C/D — the lane's MCP server is the shim: two read-only tools over ContextEngine.Search, min_score pinned to 0.15 (see §8 — do not "fix" this).
4. Test drive (Ashik's checklist)
Everything below runs on the devbox. Prereqs are already true: gwpoc.service active, dev stack up (context-engine on :8083), grants seeded.
| # | Action | Effect you should see |
|---|---|---|
| 1 | Open https://gw-dev.upsquad.ai/ui (sign in with @upsquad.ai) | the live board; possibly a small fresh capture (it was cleared 2026-08-20) |
| 2 | ANTHROPIC_BASE_URL=http://localhost:8787 ANTHROPIC_CUSTOM_HEADERS="X-Upsquad-Identity: ashik@upsquad.ai" claude -p "hello" | you appear on the board as ashik@… with true cost per call — note the hidden title-gen call and HEAD /api/hello probe (FINDINGS #7, #2) |
| 3 | Two terminals: bash /opt/upsquad/ops/gateway-poc/demo/gwdemo.sh live --profile vanilla --session ashik1 and … --profile upsquad --session ashik1 --identity ashik@upsquad.ai | two interactive Claude Codes, one per lane; paste the same prompt into both |
| 4 | Try t3: "How does a workflow approval get escalated: which RPC is used, what fields does its request take, and what happens to the node's status while it waits? Ground your answer in this project's API documentation and name the source you used." | both lanes complete (2-for-2 in rehearsal); vanilla greps (Bash×~10), upsquad retrieves (search_knowledge×2-5, usually one ⛔ denied Bash) |
| 5 | /ui?view=compare | the pair auto-selected by session name; verdict banners (✓ DELIVERED), governance chip, shared-scale token bars, delta strip, CSV export |
| 6 | In the upsquad lane, ask it to run a shell command | live platform deny quoted back verbatim (tool:claude_code.Bash — platform default deny); the ledger records the attempt |
| 7 | Ctrl+D a lane | the launcher prints that lane's ledger rollup automatically |
| 8 | (Optional) Clear capture… in the UI header | typed-CLEAR confirm; deletes ledger+sidecars from disk, counter restarts at #1, no service restart. No undo — archive first if you care |
| 9 | Unit tests: cd /opt/upsquad/ops/gateway-poc && make check | gofmt+vet+39 tests+build → builds gwpoc.new, never the live binary (see §10 cutover) |
Two disciplines while comparing: identical prompts in both lanes (live mode puts that on you), and --strict-mcp-config stays on (the launcher does this) — without it your ~47 claude.ai connector tools pollute both lanes and the A/B measures your connector list (FINDINGS #25).
5. The agent bootstrap script — demo/gwdemo.sh
One script assembles lanes, manages the shim, launches the CLI, and prints the ledger rollup. It never touches the gateway — gwpoc.service owns :8787; the launcher only reads /ui/data.
Subcommands
| Command | Signature | What it does |
|---|---|---|
run | --profile vanilla|upsquad --task <tag> --prompt "…" [--max-turns N] | one-shot claude -p in the lane, then rollup (gwdemo.sh:239-294) |
live | --profile … --session <name> [--identity you@upsquad.ai] [--dry] | interactive Claude Code in the lane; rollup on exit (:300-356) |
shim start|stop|status | — | mcpshim lifecycle (pidfile + log under demo/lanes/) |
Environment
| Var | Default | Carries |
|---|---|---|
GWDEMO_GATEWAY | http://localhost:8787 | ANTHROPIC_BASE_URL + rollup reads |
GWDEMO_SHIM_ADDR | localhost:8791 | shim -addr and the lane's .mcp.json URL |
GWDEMO_IDENTITY | vaisakh@upsquad.ai | X-Upsquad-Identity (override per-lane with --identity) |
Lane assembly (write_lane, gwdemo.sh:147-191)
| Profile | Gets | Guarantee |
|---|---|---|
upsquad | .claude/settings.json — PreToolUse matcher "Bash|Read|Glob|Grep|mcp__.*" → bash demo/hooks/govern-check.sh; .mcp.json — upsquad_context at http://localhost:8791/mcp | shim auto-started if unhealthy (prefers ./gwpoc.new mcpshim, falls back to go run . mcpshim — the service binary is never used for the shim) |
vanilla | nothing — .claude/ and .mcp.json are deleted and their absence asserted | "a vanilla lane that silently inherited a settings.json … would make the whole comparison a lie, and it would look exactly like a working demo" (:144-146) |
Both lanes always launch with --strict-mcp-config — the A/B's control arm, not hygiene (FINDINGS #25).
Session naming model (live mode)
- The
--sessionname rides theX-Upsquad-Taskheader 🟡 (one wire field, two names — you'll see it astaskin the ledger). /clearinside the CLI ⇒ a new comparison row under the same name (fresh session UUID, fresh context).- New name ⇒ exit + relaunch. The name is pinned at launch: renaming inside Claude Code never reaches the wire, and renaming mid-conversation would drag warm cache/context into the "new" group — bad science by construction.
6. Governance wiring (the upsquad lane)
The hook — demo/hooks/govern-check.sh
PreToolUse, fail-closed: any error, timeout, or unreachable engine ⇒ deny. stdin is the hook payload (.tool_name); stdout is {} (allow) or a permissionDecision: "deny" envelope with the verdict quoted (hookSpecificOutput, exit 0 either way). Targets: tool:claude_code.<Tool> for built-ins; mcp__upsquad_context__search_knowledge → tool:mcp.upsquad_context.search_knowledge (splits on the double underscore so tool names with single underscores survive). It calls the real platform PEP — plain curl to POST :8083/upsquad.governance.v1.GovernanceService/Check with X-Org-Id/X-Tenant-Id/X-Member-Id/X-Clearance — verdict parity with the original buf curl was verified before the swap. Identity keys on the triggering human (core#2564), not the agent; env-overridable (UPSQUAD_ORG/MEMBER/CLEARANCE/GOVERNANCE).
Live cascade state (verified 2026-08-21, read-only Check)
| Target | Verdict | Deciding layer | policy_id |
|---|---|---|---|
tool:claude_code.Bash | deny | platform | — (default) |
tool:claude_code.Read | allow | org_unit (GTM) | 82208236… |
tool:mcp.upsquad_context.search_knowledge | allow | org_unit | 493badad… |
tool:mcp.upsquad_context.get_context | allow | org_unit | c2c7cc06… |
Read the cascade correctly: the platform layer denies all four ("platform default deny — external egress fails closed"). The three allows win because a more-specific org_unit grant overrides the platform default. Absence of a grant is a deny — that's why the demo's denial needed no seeding, and why the grants are the interesting event.
The grant scripts 🟠
verify/grant-demo-hook-allows.sh seeds Read/Glob/Grep (floor 0, GTM unit) and deliberately not Bash; verify/grant-demo-mcp-allows.sh seeds the two shim tools and proves per-target scope (its verify shows …delete_everything in the same namespace still denied). Founder-run on purpose: "Governed means the capability does not exist until someone with authority says it does." Both verify through the same Check path the hook uses.
The honest boundary 🟡
Enforcement is client-side cooperation (FINDINGS #23): Claude Code executes tools locally, the hook is the gate, and the shim behind it has no auth of its own — curl localhost:8791 reads the corpus without any of this. The gateway observes tool intent (it appears in the ledger even when denied); it cannot block execution. Production shape = the team MCP gateway's JWT edge (ADR-0024), which exists in core but is dark on dev and not in this loop.
7. Capture, cost, verdicts (what the ledger can prove)
- Exchange row (one JSONL line, fsync'd): envelope (
n/ts/latency_ms), HTTP,key_id+auth_scheme(credential never stored — only unsaltedsha256[:12]🟡), identity + source,profile/task, client fingerprint (UA, betas), carrier-(d) decomposition (account/device/session— notedevice_idis a per-machine fingerprint, a privacy artifact, FINDINGS #1), request/response summaries,cost_usd(explicitnullwhen unknowable),maskedcount. - Attestation sidecars: masked full request/response bodies per exchange (
-full, default on). Masking is 8 deliberate rules (sk-ant, GitHub PATs/classic, AWSAKIA, PEM headers, Slackxox*, GoogleAIza, OpenAIsk-proj) — explicitly not a DLP engine (mask.go:8-11); its point is proving full-prompt capture will contain credentials. - Pricing discipline: unknown model ⇒
null, never extrapolated. A non-error response (status < 400) whose body was delivered but never parsed ⇒null/capture-blind— pricing it $0 would be confidently wrong; a failed request's zero usage is a true $0 (Anthropic doesn't bill errors). Fast mode is a request-side field — cost is not derivable from responses alone (FINDINGS #20);repricereapplies the card to history readingspeedback off the ledger.serveandrepriceshare onepriceExchangeso they cannot disagree. - Session verdicts (mechanical, from the final LLM call's
stop_reason— nothing judges):delivered(end_turn/stop_sequence) ·cut_off(tool_use — budget died mid-work, nothing shipped) ·truncated(max_tokens) ·error·none(probes only) ·unknown(an unrecognized stop_reason — the honest bucket, never coerced). Plusdenialscounted from the hook'sgovernance DENYreason in tool results 🟡 (substring-coupled to the hook's wording, 200-char sample window).llm_callsexcludes plumbing —count_tokens,HEAD /api/hello,/v1/modelsare not the denominator (FINDINGS #21). - Clear:
POST /ui/clear(GET→405), runs under the writer's lock, truncates through the held fd, resets the counter to #1, deletes sidecars, stderr audit line.identities.jsonand archives untouched. No undo.
8. The MCP shim — gwpoc mcpshim
Minimal streamable-HTTP MCP server (initialize / tools/list / tools/call, plain JSON) on 127.0.0.1:8791/mcp, GET /healthz for the launcher.
| Flag | Default | Note |
|---|---|---|
-ce | http://localhost:8083 | ContextEngine connect/grpcweb base |
-org / -member / -clearance | dev seed org / Vaisakh / 5 | 🔵 asserted, not authenticated — dev trust model; on graduation clearance must lose its default |
-min-score | 0.15 | do not "fix" this to 0 or remove it — and core#2618 landing does NOT retire it. The engine reads min_score <= 0 as "use the default 0.7" — zero is the strictest setting, not the loosest — and real corpus scores run 0.6–0.66, so the default silently returns nothing. core#2618 fixed the AssembleContext leg only, by giving it an explicit retrieval_min_score knob (default 0.3); ContextEngine.Search's wire default is deliberately unchanged, so a direct Search caller that omits min_score still gets 0.7. This shim is a direct Search caller. The pin stays. core#3079 does not retire it either — it puts the floor that ran on the response (effective_min_score, results_below_threshold, outcome), so a shim that omitted the pin would now be able to SEE that it had inherited 0.7 rather than merely receive nothing. Seeing is not sending. See docs/reference/retrieval-min-score.md §Scope. |
Tools: search_knowledge{query, top_k≤20} → per-hit [score] source — content (chunk display cap 2200 chars — raised from 700 after a mid-field cut manufactured a confidently wrong answer, FINDINGS #28: truncation in a retrieval pipe is a correctness bug); get_context{task} → grouped/deduped composition capped at 14000 chars. get_context is Search-composed, not AssembleContext — AssembleContext's own retrieval leg used to inherit the 0.7 floor with no request knob (live probe, 2026-08-21: assembledTokens: 3, fallbackReason: no_results — note that core#3079 changed what that probe would print today: an all-below-the-floor retrieval now reports below_threshold with a count, and no_results is reserved for a corpus that matched nothing. The observation stands as history; do not expect to reproduce the string). core#2618 closed that: AssembleContext now sends an explicit floor from the retrieval_min_score knob, so switching get_context back to AssembleContext is now unblocked (and is the remaining migration on this shim). The -min-score pin above is a separate matter and still stands. Errors are honest tool-results (isError: true) — "a fabricated answer built on silence is the failure mode this whole demo exists to argue against."
Known retrieval limit: single-hop recall (core#2619) — search returns a relevant chunk, sometimes not the most complete one in the same document. The grep lane caught a docs self-contradiction the RAG lane missed, both reps. Never claim answer-quality superiority.
9. What the rehearsal proved (n small, honest)
18 scripted runs (reps 1–2 + scenario B, 2026-08-17) + live pairs (2026-08-18). Provenance, precisely: the scripted runs' ledger is archived on the devbox (capture.archive-rehearsal-20260818-045242/); the headline economics line is FINDINGS.md #30 verbatim. The flagship sample1 pair's ledger rows no longer exist — they were destroyed by the operator /ui/clear of 2026-08-20 (§7's "no undo" caveat, biting its own evidence; archive before you clear). Its launcher run logs — full rollups and agent transcripts — are preserved at capture.archive-runlogs-20260821/ on the devbox, so treat sample1 as recorded-then-cleared: trust the logs, or re-run the pair (§4 step 4) for fresh disk-traceable numbers (stochastic — expect the shape, not the digits).
- Economics: governed lane ranged −13% (rep-1, derivable from the rehearsal archive) → −31% cost / −44% wall-clock (rep-2, FINDINGS #30). Variance is real; t1 flipped winners between reps.
sample1(the "blocked run" 4-part ops question): vanilla failed twice on budget (36 LLM calls, $2.01, nothing shipped) while the governed lane delivered a complete answer at $0.72 — with its Bash attempt denied by live governance (the archived transcript opens: "Bash is denied in this lane by platform governance, so everything below comes from the corpus"; the ledger's richer denial count did not survive the 2026-08-20 clear). - Containment (FINDINGS #29): "hiding" docs from the ungoverned lane failed — unrestricted Bash grepped the disk and found
.docs-api-hidden/in under a minute; the governed lane structurally couldn't. Hiding loses to an agent with a shell; governance held. - Dogfood value: the spike filed two real CE bugs with live evidence — #2618 (silent-empty) and #2619 (silent-partial) — plus 30 recorded findings (
FINDINGS.md, the half of the spike that doesn't fit in a binary — read #1–#9 before touching wire parsing).
10. Ops
| Concern | Fact |
|---|---|
| Service | /etc/systemd/system/gwpoc.service → ExecStart=…/gwpoc serve -addr localhost:8787, User=vb, Restart=always, loopback-only |
| Exposure | CF tunnel ingress gw-dev.upsquad.ai → localhost:8787; CF Access app allows @upsquad.ai; kill-switch sudo bash teardown-gw-exposure.sh (Access → DNS → ingress → unit, idempotent) |
| Cutover discipline | make build → gwpoc.new; make build-live is the explicit opt-in that overwrites ./gwpoc (the path systemd execs). This exists because make check once hot-swapped the live binary and nothing broke — which was the problem (FINDINGS #26). Cutover = stop unit → swap → start; gwpoc.prev is the rollback |
| Reprice | ./gwpoc.new reprice -capture capture/exchanges.jsonl [-dry] — writes .bak (single-level 🟡), atomic rename |
| Data custody | capture*/ dirs are devbox-only forever (real prompt attestation = evidence, not code; purged from git history before first push, gitignored as a class). verify/demo.sh refuses to run against the live unit and 2 of its 7 legacy assertions fail by design (FINDINGS #27 — the code improved past them) |
11. Gaps / dev-only / graduation gates
If the spike graduates, these are the gates (also on record in the client#804 review):
- 🔵 Shim auth — none today; the dev-stack ICS trust model (
DISABLE_AUTH=true+ICS_MODE=trueon memory-mcp is the same pattern). Prod = the team-gateway JWT edge. - 🟡 Enforcement locus — PreToolUse is cooperative, client-side. Real containment needs the network boundary.
- 🟠 OAuth ToS — subscription auth transiting our infra (FINDINGS #3, CAVEATS #2) is a Terms-of-Service question, founder call, before anything customer-facing.
- 🟠 Capture custody ADR — full-prompt capture (~260 KB/exchange incl. file contents) is a data-governance event, not a feature toggle (CAVEATS #4). Blind spots: the final turn's tool results and compaction (CAVEATS #5).
- 🟡 Rate card is a snapshot — re-verify against the pricing page before quoting; 1-hour cache writes, data-residency 1.1×, Batch 0.5× unmodelled.
- 🟡 Known code quirks a tester may hit:
/uiendpoints have no auth of their own (loopback + CF Access are the fences);response.bytesis wire bytes, not decoded; request bodies buffer uncapped in RAM;br/zstdresponses get norespsidecar;identities.jsonloads once (no reload); README in-tree still says 6000-char context cap and "~400-line shim" (stale — code is 14000 and 600; this doc is current). - 🟠 The direction itself — core#2616 pending. This doc describes evidence, not a decision.
12. Source map
Path (under tools/gateway-poc/) | Owns |
|---|---|
serve.go | proxy pipeline, streaming, header hygiene, fail path |
identity.go | 4 carriers, key fingerprinting, user_id parsing |
capture.go / mask.go | ledger schema, sidecars, Clear, the 8 mask rules |
pricing.go / reprice.go | rate cards, priceExchange, history reprice |
sessions.go | session rollup, outcomes, denial counts |
ui.go / ui.html | /ui, /ui/data, /ui/body, /ui/clear, Compare view |
mcpshim.go | the upsquad_context MCP server |
demo/gwdemo.sh | the agent bootstrap script (lanes, shim, run/live) |
demo/hooks/govern-check.sh | the fail-closed governance hook |
verify/ | no-key demo harness, founder grant scripts, fixtures |
report.go, fakeupstream.go | CLI report; deterministic fake api.anthropic.com |
FINDINGS.md | 30 findings — the spike's real deliverable alongside the binary |
Doc reflects upsquad-client@49ca8e4e (gwpoc v0.6.0) + live devbox state verified 2026-08-21. Invalidated by: a new gwpoc version, changes to the governance policy seeds, the #2619 fix, or a decision on core#2616. #2618 landed and this page was updated for it: it unblocked the shim's get_context Search-composition, and it did not retire the -min-score pin — see the two rows above. core#3079 landed and this page was updated for it too: the retrieval contract now names below_threshold separately from no_results, which changes the string a re-run of the §8 probe prints, and it likewise does not retire the pin.