Skip to main content

ML Detector API

Table of Contents​

Top

upsquad/ml/v1/detector.proto​

ClassifyRequest​

ClassifyRequest carries the text to score plus attribution metadata used only for metric labels and structured logs.

FieldTypeLabelDescription
textstringtext is the user message to classify. Caller MUST enforce a 32KB soft ceiling; the service rejects >64KB with INVALID_ARGUMENT and logs 32–64KB as oversized-but-truncated.
tenant_idstringtenant_id is the calling organization UUID. Used for metric labels only; the sidecar does not authorise per-tenant access.
session_idstringsession_id is the agent session UUID. Used for log correlation.
agent_idstringagent_id is the invoking agent UUID. Used for metric labels.
model_hintstringmodel_hint selects the classifier model. Wave 5 accepts exactly "prompt-guard-86m"; any other value returns INVALID_ARGUMENT.

ClassifyResponse​

ClassifyResponse carries the classifier verdict plus raw scores so the orchestrator can emit shadow-phase metrics without re-thresholding.

FieldTypeLabelDescription
jailbreak_scorefloatjailbreak_score is the Prompt-Guard JAILBREAK class probability, range [0.0, 1.0].
indirect_injection_scorefloatindirect_injection_score is the Prompt-Guard INJECTION class probability, range [0.0, 1.0]. Reserved for Wave 5.1 tuning.
verdictstringverdict is one of "allow", "warn", "block" — computed server-side from the tenant-provided thresholds passed in ClassifyRequest.
latency_msint32latency_ms is the sidecar-side wall-clock inference time in ms.
model_versionstringmodel_version identifies the classifier build for audit rows, e.g. "prompt-guard-86m@v1".
classified_atgoogle.protobuf.Timestampclassified_at is when the sidecar finished inference.

DetectPIIRequest​

DetectPIIRequest carries one free-text fragment to scan (#2155).

The unit is a fragment, not a document: the Go seam calls Scrub(s string) on every string LEAF of a memory candidate's JSON content, so the caller has already decomposed the payload.

FieldTypeLabelDescription
textstringtext is the fragment to scan. Same 64KB ceiling as ClassifyRequest; over that the service returns INVALID_ARGUMENT.
tenant_idstringtenant_id is the calling organization UUID. Metric labels only.
session_idstringsession_id is the agent session UUID. Log correlation only.
agent_idstringagent_id is the invoking agent UUID. Metric labels only.
model_hintstringmodel_hint selects the PII model. Accepts exactly "gliner-small-v2.1"; any other value returns INVALID_ARGUMENT. The gate exists so that a config drift to a different artifact fails loudly at the first call rather than silently serving predictions from an unevaluated model.

DetectPIIResponse​

DetectPIIResponse carries the located spans plus the evidence a caller needs to decide whether to TRUST them.

FieldTypeLabelDescription
spansPIISpanrepeatedspans are the accepted PII spans, ordered by start offset. Empty means the fragment is clean — but ONLY when canary_healthy is true.
latency_msint32latency_ms is sidecar-side wall-clock inference time.
model_versionstringmodel_version identifies the artifact, e.g. "urchade/gliner_small-v2.1@<revision-sha>".
detected_atgoogle.protobuf.Timestampdetected_at is when the sidecar finished inference.
canary_healthyboolcanary_healthy reports whether the ANTI-COLLAPSE canary is passing.

This field exists because of a measured failure mode, not a hypothetical one. The int8-quantized build of this exact model returns ZERO detections on the entire corpus while loading cleanly, reporting SERVING, passing the readiness probe, and running 35% FASTER (#2155 §5). Every signal the platform had said "healthy"; the detector was inert. A silently-empty span list is therefore indistinguishable from a clean fragment unless the server proves, per response, that it can still detect a sentence it is KNOWN to detect.

When false the server returns UNAVAILABLE rather than a response, so a caller that never reads this field still fails closed. It is carried anyway so the failure is legible in a grpcurl probe. |

ModerateRequest​

ModerateRequest is the LLD 23 placeholder. Wave 5 LLD 22 returns UNIMPLEMENTED from the server; shape may evolve in LLD 23 without a breaking change by extending reserved field numbers.

FieldTypeLabelDescription
textstringtext is the content to moderate.
tenant_idstringtenant_id is the calling organization UUID.
session_idstringsession_id is the agent session UUID.
agent_idstringagent_id is the invoking agent UUID.

ModerateResponse​

ModerateResponse is the LLD 23 placeholder. Wave 5 LLD 22 returns UNIMPLEMENTED; LLD 23 defines the fields.

PIISpan​

PIISpan is one located span of PII in the request text.

FieldTypeLabelDescription
pattern_classstringpattern_class is a member of the BOUNDED platform enum, never a raw model label. GLiNER is zero-shot, so its label set is an INPUT — the service maps its own prompt labels onto this enum and drops anything unmapped, which makes the enum closed BY CONSTRUCTION rather than by a hand-maintained mapping table that a future upstream release lands outside of.

Emitted values are exactly: "person", "dob", "address".

natid is DELIBERATELY ABSENT. The prompt still asks for national-ID labels — removing them would change the zero-shot label interaction and therefore invalidate the measured configuration — but the predictions are discarded server-side, because the Go caller derives that class from format + checksum instead. See the DetectPII doc. | | start | int32 | | start is the UTF-8 BYTE offset of the first byte of the span, half-open with end.

BYTES, not codepoints, and the distinction is load-bearing rather than pedantic: the server is Python (which indexes str by codepoint) and the caller is Go (which indexes string by byte). The #2155 corpus contains Aoife Ní Bhriain and 78 Rue de Rivoli — multi-byte runes sitting to the LEFT of a gold span — so a codepoint offset handed to Go slices the wrong characters and redacts the wrong text. The server converts before it serialises. | | end | int32 | | end is the UTF-8 byte offset one past the last byte of the span. | | score | float | | score is the model's confidence in [0.0, 1.0]. Spans below the server's threshold are already dropped; this is carried for tuning telemetry. |

MLDetectorService​

MLDetectorService hosts the Wave 5 ML inference RPCs. Runs in the ml-detector sidecar Deployment (namespace: platform). All RPCs are UNARY. The Go bridge in internal/runtime/security/mlclassifier calls Classify; LLD 23 fills the Moderate stub.

Method NameRequest TypeResponse TypeDescription
ClassifyClassifyRequestClassifyResponseClassify scores a single user message for prompt-injection / jailbreak risk using Prompt-Guard-86M. Latency budget: P95 <75ms including one RTT from the orchestrator. The service must fail fast on any input larger than 64KB with INVALID_ARGUMENT.
ModerateModerateRequestModerateResponseModerate scores a text buffer for policy-category violations. LLD 22 returns UNIMPLEMENTED; LLD 23 fills the body. Request and response shapes are minimal here to let LLD 23 land without a second proto.
DetectPIIDetectPIIRequestDetectPIIResponseDetectPII locates UNSTRUCTURED PII spans in a free-text fragment using GLiNER (urchade/gliner_small-v2.1, zero-shot), for the memory-extraction redaction seam (#2155).

It answers exactly ONE half of the hybrid detector. The model owns the classes that have no parseable form — person, dob, address — and the caller owns natid, which is deterministic format + checksum and is computed in Go. That split is not tidiness: measured on the #2155 corpus, the model reaches 0.62 strict recall on national IDs while the format list reaches 0.92 with ZERO false positives on all 65 engineering negatives. The service therefore NEVER returns natid; see the note on PIISpan.pattern_class.

Latency budget: p95 <250ms (measured 119ms on 8 vCPU CPU-only). This sits on the ASYNC extraction path, not on a request path, so the budget is loose by design.

Fails with UNAVAILABLE when the anti-collapse canary is failing — see DetectPIIResponse.canary_healthy. Callers MUST treat any error as "detector unavailable" and fall back to regex + quarantine-all rather than to no detection (#2155 scope item 4). |

Scalar Value Types​

.proto TypeNotesC++JavaPythonGoC#PHPRuby
doubledoubledoublefloatfloat64doublefloatFloat
floatfloatfloatfloatfloat32floatfloatFloat
int32Uses variable-length encoding. Inefficient for encoding negative numbers – if your field is likely to have negative values, use sint32 instead.int32intintint32intintegerBignum or Fixnum (as required)
int64Uses variable-length encoding. Inefficient for encoding negative numbers – if your field is likely to have negative values, use sint64 instead.int64longint/longint64longinteger/stringBignum
uint32Uses variable-length encoding.uint32intint/longuint32uintintegerBignum or Fixnum (as required)
uint64Uses variable-length encoding.uint64longint/longuint64ulonginteger/stringBignum or Fixnum (as required)
sint32Uses variable-length encoding. Signed int value. These more efficiently encode negative numbers than regular int32s.int32intintint32intintegerBignum or Fixnum (as required)
sint64Uses variable-length encoding. Signed int value. These more efficiently encode negative numbers than regular int64s.int64longint/longint64longinteger/stringBignum
fixed32Always four bytes. More efficient than uint32 if values are often greater than 2^28.uint32intintuint32uintintegerBignum or Fixnum (as required)
fixed64Always eight bytes. More efficient than uint64 if values are often greater than 2^56.uint64longint/longuint64ulonginteger/stringBignum
sfixed32Always four bytes.int32intintint32intintegerBignum or Fixnum (as required)
sfixed64Always eight bytes.int64longint/longint64longinteger/stringBignum
boolboolbooleanbooleanboolboolbooleanTrueClass/FalseClass
stringA string must always contain UTF-8 encoded or 7-bit ASCII text.stringStringstr/unicodestringstringstringString (UTF-8)
bytesMay contain any arbitrary sequence of bytes.stringByteStringstr[]byteByteStringstringString (ASCII-8BIT)