Knowledge registry backfill reports (M4)
cmd/knowledge-backfill writes one YYYY-MM-DD.md here per run. This is where
M4 — "every existing chunk resolves to exactly one registry row" is measured.
T4 (#3076) · LLD #3072 §3.5 · HLD #3046 §3.5 · PRD #2995 TK-2.6 · tracker #2991.
Running it
# Preview. Runs the real passes on a real transaction and rolls back, so the
# numbers are exactly what --apply would produce.
DATABASE_URL=... go run ./cmd/knowledge-backfill
# Apply, one org at a time or all of them.
DATABASE_URL=... go run ./cmd/knowledge-backfill --apply [--org <uuid>]
Exit codes: 0 M4 holds · 1 M4 is red (the report names every offending
address) · 2 infrastructure or usage error.
The org set is the outer denominator, and an empty one is refused. A --org
naming no organisation, and a DSN whose database holds none, both exit 2 and
write no report — a run that measured nothing must not publish a verdict.
The verdict line also distinguishes a pass over a populated corpus from a pass
over an empty one, because "every one of the 0 addresses resolves" is a sentence
a reader cannot tell apart from a typo.
A dry run will not overwrite the same day's APPLIED report. The filename is
the date and carries no mode marker, so a preview after an apply would rewrite
the audit entry underneath a reader. Pass --date to write elsewhere.
Reading the report
The report is a snapshot and that is why it is dated. The legacy population shrinks as sources sync, so a number here is true of the day in the filename and of no other day.
Four buckets partition the legacy population — every knowledge_doc_versions
row with raw_sha256 IS NULL — and they must sum to it. A run whose buckets
do not sum exits 2 rather than publishing.
| bucket | meaning | merged by |
|---|---|---|
merged_by_content_address | upload/web legacy row, bytes unchanged since the backfill. Merged, not repaired — see below | the ingest writer, DocArtefacts.Record |
merged_by_legacy_gdrive_address | gdrive legacy row adopted onto its real document | this command |
legacy_content_changed | upload/web only. The source ingested real content after this row and this row was not it — an upper bound on divergence, not a per-document fact. See below | — |
unmerged_pending_resync | the source has not synced yet — a state, not a failure | — |
The buckets are derived from END STATE rather than from a merge ledger, so they are correct whichever of the two writers performed the merge.
Why two writers, and why that is not a split brain. The merge key set is two forward derivations (#3076 amendment §3), and each is computable in exactly one place:
GdriveLegacyVersionAddress(source_id, drive_file_id)— the operand is persisted asknowledge_docs.external_id, so this command derives it from the registry alone and converges gdrive legacy rows on every run;SourceDocIDForContent(source_id, sha256(raw))— the operand exists only while the raw bytes do, which is the ingest path at the moment of ingest. Nothing can recover it afterwards: a legacy row has no hash by construction.
Both go through one store primitive, Store.AdoptLegacyDocVersion, so the
four-step order the schema forces cannot be right in one place and wrong in the
other. There is no sequencing constraint between this command and the first
post-repair sync — both orders converge on one document (amendment §4b).
A merged_by_content_address row is merged, not repaired
Read this before treating the bucket as "done". The content-key merge
re-parents the version and returns before the artefacts are persisted, so the
surviving document holds one version with raw_sha256 NULL,
artefact_state='raw_unavailable', and no blob refs — permanently:
- the bytes were in the crawler's hand at the moment of the merge and are discarded there;
- the next unchanged re-crawl finds the address resolved (now to this same document) and short-circuits again;
- #3094's promotion pass cannot reach it — its population is
raw_unavailablewith a hash, and there is no hash here to promote.
So after KNOWLEDGE_ARTEFACT_BUCKET is wired in prod, this subset still
archives nothing. Counting the row as merged is correct; reading it as a
complete document is not. The population is raw_sha256 IS NULL and the
document no longer carrying legacy:.
The repair needs its own design and its own pass — it is a fourth condition on
#3094, recorded in #3076 ruling (3) §3. It is deliberately not fixed inside
the merge: filling the row in there would move it out of raw_sha256 IS NULL
and make merged_by_content_address structurally unreachable, and it would
write on unchanged content — the property #3074's idempotency proof pins.
Which sources can actually reach merged_by_content_address
The bucket is reachable only where an ingest of unchanged bytes actually calls the registry writer. Measured per path:
| path | unchanged-content guard | reaches the merge? |
|---|---|---|
upload / paste (ingest.go) | none — always calls Record | yes |
gdrive (gdrive.go) | exists && storedHash == contentHash → continue | no |
web (connector.go) | res.NotModified || (res.ContentHash != "" && res.ContentHash == prevHash) → unchanged | no — see the precondition |
The web guard compares the body hash persisted by RecordWebRefresh, so it
is ETag-independent and fires on every unchanged re-crawl. Changed bytes derive
a different address, so there is no legacy row there to merge either.
The precondition, named rather than implied. That short-circuit needs
prevHash to be non-empty — content_hash is nullable, not NOT NULL, and
the guard cannot fire on NULL. It is nevertheless true for every web source
that has ever ingested, because the web connector and the content_hash
column both arrive in migration 106: no web source can hold chunks written
before the column existed. So the claim rests on migration history, not on
the schema, and that is worth saying out loud — a schema-only reading of it is
false.
Consequence for a web-only tenant: merged_by_content_address cannot be
non-zero, and a legacy web row's terminal state is unmerged_pending_resync
(or legacy_content_changed) — never merged. That is a fact about what the
number means for them, not a defect.
The bucket is NOT renamed, and that is a ruling. It is named for the key that performed the merge, which is what it counts and stays true for every path. A path-flavoured name would bake in a fact about today's guards — remove the web unchanged-guard, or add a path without one, and the name becomes false while the code stays right. Reachability is a property of the callers, which is why it lives in this table. Raised by QA on PR #3105 (F4); ruled by the architect.
legacy_content_changed is an upper bound, and never holds a gdrive row
For gdrive the answer is per-document and provable: the adopt pass runs
before the measurement and adopts any legacy row whose file already has a
document, so a gdrive row still unmerged at measurement time has no target yet.
It is unmerged_pending_resync however many sibling files in the same folder
have synced.
For upload and web there is no per-document link to have — a changed document derives a different address and nothing joins the two — so the signal is source-scoped: a sibling document's ingest moves every remaining legacy row in that source into this bucket. Read it as "the source moved on without this row", not as "these bytes diverged".
Two populations that are counted and excluded, on purpose
- Conversation chunks (
rag_chunks.source_id IS NULL, migration 099) are not documents:knowledge_docs.source_idisNOT NULLand referencesknowledge_sources. They are reported so M4's denominator is auditable, and excluded from it because a metric measured over data nobody can act on is red forever. - Chunks under a soft-deleted source cannot be given a registry row and
fail the report by name.
DeleteSourcehard-deletes chunks (founder D2), so surviving chunks under a deleted source are drift.
Every backfilled row is raw_unavailable, and that is the answer
Nothing ever persisted the legacy corpus's bytes. source_doc_id is a one-way
UUIDv5 and rag_chunks.content_hash is per-chunk, so the document's raw hash was
never stored anywhere. raw_sha256, content_type and size_bytes are
therefore NULL, discriminated by artefact_state — a sentinel was rejected
(ruling on #3076), because legacy:unknown in a hash column is exactly the
"populated while holding nothing" state migration 228's CHECKs exist to forbid.
TK-2.6 asks for precisely this: recorded, never silently omitted.