Runbook — memory-extraction NER redaction detector (#2155)
Scope. The PII detector behind the memory-extraction redaction seam:
internal/context/memory/redact in context-engine, plus the DetectPII RPC
on the ml-detector sidecar.
Prod status. MEMORY_EXTRACTION_ENABLED is OFF in production. Turning it
on is a separate founder decision; this detector landing is a precondition for
that decision, not the decision itself.
1. What this thing actually protects against, and what it does not
Memory extraction mines session transcripts into durable memories that are recalled into agents' context later. That makes a transcript's PII a leak vector with a long half-life. The detector runs on every candidate before it is persisted.
It is a hybrid, split by how each class is derived:
| class | derived by | tier | on a hit |
|---|---|---|---|
person, dob, address | GLiNER model in the sidecar | LOW | flag + quarantine, text left verbatim |
natid | deterministic format + checksum, in Go | HIGH | rewritten to [redacted:natid] + quarantine |
secret, ssn, cc, email | regex denylist | HIGH | rewritten + quarantine |
phone, ip | regex denylist | LOW | flag + quarantine, verbatim |
detector_unavailable | sentinel — not a detection | LOW | quarantine only |
LOW tier means the reviewer sees the original bytes. That is deliberate: a
reviewer shown 170[redacted:phone] cannot tell a leaked phone number from a
Unix timestamp, and there is no pre-redaction original stored anywhere.
Known gaps — read these before trusting a green dashboard
- ENGLISH ONLY. Non-English names, dates of birth and addresses are not detected and pass through. Accepted v1 scope (#2155). If an org's sessions are not in English, this detector is not protecting them.
natidis a denylist of known jurisdictions (IE, UK, IN, CA, ES, IT, passports, UK driving licence). Anything else is not detected.personfalse positives on engineering eponyms (Luhn,Alice,Envoy,Martin Fowler) — roughly 6–8% of engineering-content records. Non-destructive by design; the cost is review-queue load.
2. Alerts and what they mean
ml_detector_pii_canary_failures_total > 0 — the important one
Meaning: the model is loaded but INERT. It is not detecting anything (or is detecting everything). PII is not being found.
This alert exists because a collapsed model is invisible everywhere else: the
int8 build of this artifact loads cleanly, reports SERVING, passes readiness,
and runs faster while returning zero detections. Health and latency cannot
tell you the detector works. Only the canary can.
What happens automatically: the DetectPII health key goes NOT_SERVING,
the RPC returns UNAVAILABLE, and context-engine falls back to regex +
quarantine-all. You are not leaking; you are flooding the review queue.
Triage:
# 1. Which direction did it collapse?
kubectl logs -n platform deploy/ml-detector | grep -E "pii.canary"
# pii.canary.collapsed -> detects nothing (the int8 shape)
# pii.canary.over_detecting -> detects PII in clean prose
# 2. Is the baked artifact the approved one?
kubectl exec -n platform deploy/ml-detector -- cat /opt/ml-detector/pii-model/PINNED_REVISION
# expect: urchade/gliner_small-v2.1@4e091416cf7c3481db542c2a3d26156916f3a47f
# 1d4e83e4e4ae4ae0a4fbc81a32ee6de480fb341650d73e808088bb2800312de4
# 3. Probe it directly.
grpcurl -plaintext -d '{"text":"Spoke with John Smith (DOB 12 March 1988, 14 Elm Street, Dublin).","model_hint":"gliner-small-v2.1"}' \
ml-detector:50060 upsquad.ml.v1.MLDetectorService/DetectPII
Most likely causes, in order: the image was rebuilt with a quantized dtype; the model volume was remounted with different weights; the pod is under memory pressure and torch is degrading.
Do not "fix" this by relaxing the canary. A detector that cannot fail visibly is the defect class this whole change exists to remove.
memory_redaction_hits_total{pattern_class="detector_unavailable"} rising
The sidecar is unreachable from context-engine. Redaction is still running
(regex + natid) and every candidate is quarantining.
kubectl get pods -n platform -l app=ml-detector
kubectl exec -n platform deploy/context-engine -- env | grep ML_DETECTOR_ADDRESS
grpc_health_probe -addr ml-detector:50060 -service upsquad.ml.v1.MLDetectorService.DetectPII
Safe to leave running: this is the fail-closed path working as designed. It is not safe to leave forever — an unworkable review queue trains rubber-stamping, and a rubber-stamped queue measures nothing.
memory_redaction_hits_total{pattern_class="other"} > 0
A registration bug, not traffic. Some class is being emitted that
memmetrics has not been taught about, so its hits are folding into one bucket
and are invisible in every per-class query. Nothing emits other in a correct
build. See internal/context/memory/redact/classes_test.go.
Review queue growing faster than it is worked
sum by (pattern_class, tier) (rate(memory_redaction_hits_total[1h]))
A rising low share is the signal to tune the detector, not to loosen the
review gate.
3. Configuration
| env var | where | default | notes |
|---|---|---|---|
REDACT_DETECTOR | context-engine | regex | ner enables this detector. regex is fail-open for names/DOBs/addresses. |
ML_DETECTOR_ADDRESS | context-engine | ml-detector.platform.svc.cluster.local:50060 | shared with the orchestrator's classifier client |
ML_DETECTOR_PII_ENABLED | ml-detector | true | false makes DetectPII UNIMPLEMENTED, which the Go side treats as an outage |
ML_DETECTOR_PII_CANARY_INTERVAL_SECONDS | ml-detector | 60 | |
ML_DETECTOR_USE_FAKE_PII_DETECTOR | ml-detector | false | never in production — reports model_version=fake-pii-detector |
There is no configuration that yields "no detection and no quarantine." Every failure path degrades to regex + quarantine-all. If you think you have found one, that is a bug worth paging someone about.
Selecting REDACT_DETECTOR=ner without a reachable sidecar logs at ERROR and
falls back to regex — check for
redact: NER detector construction failed at context-engine startup.
4. Verifying the detector actually works (not just that it is green)
# Fast, model-free: the deterministic half + the registration guard.
go test ./internal/context/memory/redact/... ./internal/context/memory/memmetrics/...
# The real artifact against the frozen corpus, with floors asserted.
python3 services/ml-detector/eval/score_corpus.py \
--model-path /opt/ml-detector/pii-model --assert-floors
The second is the one that matters after any model change. Green health plus green unit tests does not establish that the model detects anything — that is precisely the int8 lesson.
5. Escalation
| situation | who |
|---|---|
| canary failing, artifact digest correct | principal-architect (model evaluation owner) |
| digest mismatch / image rebuild needed | devops-engineer |
| changing the model identity | founder — new dependency decision |
| re-tiering any class | founder — at-rest security posture (ADR-0029 D3) |
enabling MEMORY_EXTRACTION_ENABLED in prod | founder |
References
- #2155 (this work), #2134 (the seam), #2095, #2499 (tiering), #2508 (withhold gate)
- ADR-0029 D1 (artifact pinning), D3 (at-rest trade), D5.1 (fail-closed tiers), OQ-3 (refresh cadence)
services/ml-detector/README.md— model, canary, air-gap detail