Skip to main content

MG-4.2 — the Model Gateway migration baseline

What it is: cmd/mg-baseline-report scans llm_usage_effective_cost over a window and reports how much of the agent fleet actually routes through the Model Gateway, per cohort and per agent, plus a latency comparison between the two cohorts.

What it is for: it is the instrument MG-4.2 asks for, and the input the MG-4.3 default flip (task 4F, #3148) is decided against.

Task 4E of LLD #3138 · issue #3147 · PRD #2644 MG-4.2 / MG-NFR.8 · milestone 31.


1. Read this before you read a number​

The flag is not the oracle​

agent_configurations.gateway_routed (migration 240) is intent. Every migration number here comes from the ledger (#3060 constraint 2).

That distinction has teeth since task 4B (#3144). Flag-on-with-no-traffic is a real, reachable state: an agent whose route minter is unarmed (MODEL_GATEWAY_BASE_URL unset), or an agent with no team, refuses the turn rather than falling back to a direct dial. So gateway_routed = true never silently means "went direct" — but it very often means "went nowhere". Counting the flag would score every one of those as migrated.

The flag is published once, as mg_baseline_agents_flag_on_diagnostic. A gap between it and mg_baseline_agents_observed_routing_measured is the count of agents whose turns are failing.

The availability half is assumed_unknown, and that is not a TODO​

MG-NFR.8 asks whether the gateway is at least as available as the worker-direct path. The ledger cannot answer it, for two structural reasons:

  1. A call that produced no row is invisible in both cohorts. A total gateway outage writes nothing. Zero rows is not 100% availability.
  2. The failure vocabulary is unstorable on the direct cohort: migration 219's termination_only_on_allow permits a termination reason only beside governance_outcome = 'allow', and migration 203 leaves governance_outcome NULL on every worker-side row. upstream_stalled cannot exist on a direct-path row even in principle.

So this tool publishes a volume and latency comparison and refuses to name it an availability one. Closing MG-NFR.8's availability clause needs a signal from outside the ledger and is an explicit input to the 4F decision — not a number this tool will invent.

A routed call writes TWO rows​

internal/modelgateway/recorder.go writes the governed row; the worker's own MetricsEvent still reaches internal/runtime/streaming/handler.go, which records an unassessed row for the same call. Therefore:

  • gateway_rows / all_rows is capped near 50% by construction — never use it.
  • The migration measure is agent-grain: an agent seen on the gateway at all is observed routing, and the shadow row lands under the same agent_id.
  • cohort="unassessed" is a row count, not a direct-call count. The direct-call reading is published as direct_calls_upper_bound and is _modelled.

Honesty suffixes​

suffixmeaning
_measuredcounted from ledger observations
_modelledderived from measured inputs by a rule stated at the field
_diagnosticderived from configuration (the flag), never from traffic
assumed_unknownnot derivable from this data at all — never a number

2. Running it​

# One-shot against the dev database, JSON to stdout.
DATABASE_URL=postgres://upsquad:upsquad@127.0.0.1:5432/upsquad?sslmode=disable \
go run ./cmd/mg-baseline-report --window=720h --json=-

# The committed audit-trail form.
DATABASE_URL=... go run ./cmd/mg-baseline-report --window=24h \
--markdown-dir=docs/runbooks/mg-baseline-reports

Exit codes:

codemeaning
0a measurement over a non-empty denominator
1insufficient_data — the window measured nothing
2infrastructure fault (no DSN, DB unreachable, unwritable output)

An empty window exits 1. There is deliberately no synthetic "baseline mode" that stamps a clean report when DATABASE_URL is unset: a scheduler, a CI step and a human all read exit 0 as "fine".


3. The dev deployment, and its two manual steps​

The compose service mg-baseline-report rescans every 15 minutes and writes a node_exporter textfile. The series arrive under the existing node scrape job — no prometheus.yml change and no scrape-config reload.

Neither of these happens on merge. Run both on the devbox:

cd /opt/upsquad/upsquad-core && git pull

# 1. node-exporter must be RECREATED to pick up --collector.textfile.directory.
docker compose -f docker-compose.dev.yml up -d --force-recreate node-exporter

# 2. start the reporter.
docker compose -f docker-compose.dev.yml up -d --build mg-baseline-report

# 3. alert RULES do need a reload — nothing on the devbox does it (#3205).
curl -sSf -X POST http://127.0.0.1:9092/-/reload
curl -s 'http://127.0.0.1:9092/api/v1/rules' | \
jq '[.data.groups[] | select(.name=="mg-baseline.alerts") | .rules[].name]'

Verify the series exist:

docker exec upsquad-node-exporter ls -l /var/lib/node_exporter/textfile
curl -s 'http://127.0.0.1:9092/api/v1/query?query=mg_baseline_last_run_timestamp_seconds' | jq .

MGBaselineInstrumentAbsent firing after the merge means these steps have not been run. That is the alert working.

Updating the reporter​

It is published to GHCR as upsquad-core-mg-baseline-report (publish-images.yml leg + image: in compose), so scripts/dev-reconcile.sh sees it by digest like every other service and a merged change reaches the running container.

An earlier revision of this PR shipped it build:-only and documented that as a known limitation. That was wrong, and scripts/check-compose-publish-coverage.py is a hard gate precisely so it cannot be documented away: a service the reconciler cannot see is indistinguishable from one that is up to date, and a stale reporter keeps refreshing its textfile happily while MGBaselineReportStale stays quiet. That is the same absence-scored-as-a-pass shape this reporter exists to detect.

To rebuild locally without waiting for GHCR:

docker compose -f docker-compose.dev.yml up -d --build mg-baseline-report

4. Known state of the dev fleet (measured 2026-09-07)​

The first live run against the devbox found, over a 30-day window:

  • 0 ledger rows across 2 orgs ⇒ verdict insufficient_data, exit 1. The dev fleet has emitted no LLM usage rows at all in that window, and the instrument refuses to call that clean.
  • agent_configurations.gateway_routed does not exist on the devbox database. Migrations 4A/4B are merged; that database has not been migrated to 240. Both orgs' flag diagnostic errored and the scan was partial — which is why mg_baseline_orgs_failed_measured exists and is alerted on.

MG-4.2 cannot be discharged from this state. Establishing the baseline needs the dev database migrated, agents opted in, and real traffic — in that order.


5. Troubleshooting​

SymptomRead
MGBaselineInstrumentAbsent§3 — the two manual deploy steps
MGBaselineReportStalethe textfile froze; every scraped value is a snapshot. Do not decide on it
MGBaselineScanPartialthe schema version first (gateway_routed missing is the case that has occurred), then RLS role, then a genuine bucket-partition fault
MGBaselineCohortSplitUntrustworthya producer is writing gateway-shaped ledger rows the reporter does not model. Code bug — do not silence
share series missing from Prometheuscorrect when the denominator is zero. 0/0 is not 0%
a cohort's p99_of_measured missingcorrect when that cohort has no measured sample, and when a fleet scan spans more than one contributing org (a percentile of a union is not a function of per-org percentiles)