Embedding footprint — measured results
NFR-3. LLD #2714 §5 T14 · PRD #996 NFR-3 / FR-19 · issue #2736 · milestone 32.
These are measurements, not prose. PRD #996 FR-19 states it in as many words —
"estimates are not acceptable in place of measurements" — so every number
below came out of scripts/embedding-footprint-bench.py on a CPU-only host, and
the raw output is in the JSON beside this file.
The operator-facing rendering of these numbers is the embedding configuration reference. This directory is the evidence; that page is where an operator chooses. Sizing decisions should cite this file.
Measurement host and conditions
| CPU | AMD EPYC-Genoa, 8 vCPU |
| RAM | 15 GiB total |
| GPU | none — no /dev/nvidia*, and Ollama reports library=cpu, layers.offload=0, size_vram=0 on both models |
| Server | ollama/ollama:0.6.8 (the digest the dev stack runs) |
| Wire shape | POST {base_url}/v1/embeddings — the OpenAI-compatible shape, which is the one the shipped embedder uses |
| Fixture | upsquad-platform-v1 — the same 40 documents / 20 queries the recall benchmark scores |
| Date | 2026-09-02 |
The fixture is shared with the recall benchmark on purpose: a throughput number taken on text we do not embed is a number about nothing in particular. Documents are 378 characters at the median, queries 86.5 — which is why the two throughput figures below are reported separately rather than averaged.
size_vram == 0 on both models is the assertion that makes "CPU-only" a
measurement rather than a claim about this host. Verify it before trusting any
re-run: a number taken with a GPU quietly present is not comparable to these.
Results
| shipped default | quality tier (operator-supplied) | |
|---|---|---|
| Model | nomic-embed-text:v1.5 | gte-Qwen2-1.5B-instruct |
| Artifact measured | r=970aa74c0a90 (the shipped pin) | hf.co/second-state/gte-Qwen2-1.5B-instruct-GGUF:Q4_K_M |
| Parameters | 137M | 1.78B (as the GGUF reports it) |
| Dimensions | 768 | 1536 (native) |
| Disk | 274.3 MB | 1116.7 MB (4.07×) |
| RAM resident at inference | 370.0 MB | 1566.7 MB (4.23×) |
| Bulk ingest | 11.5 embeddings/sec (10.4–12.5) | 2.0 embeddings/sec (2.0–2.0) — 5.65× slower |
| Query path p50 | 32.3 ms | 171.3 ms (5.31×) |
| Query path p95 | 46.1 ms | 235.4 ms |
Raw output: footprint-local-nomic.json ·
footprint-quality-tier-gte-qwen2.json.
The Dimensions row is a spec, not a result — and the cost rows below it are not caused by it. This table has a 768 column and a 1536 column stacked above four rows the 768-d model wins by 4–5×. That is a comparison of two specific artifacts — a 137M-parameter f16 model against a 1.78B-parameter Q4_K_M requantization — and the cost gap tracks the 12.99× parameter count, not the 2× vector width. This milestone measured no dimensionality effect and licenses no claim about one, in either direction. A hypothetical 1536-d model at 137M parameters is not measured here and would not be predicted by this table. Retrieval quality of the 1536-d column was not measured at all — see below.
What each number is, exactly
- Disk is what the model store holds — Ollama's
/api/tagssize, which is the manifest's config plus every layer (weights + params + license), not the weights blob alone. For the shipped artifact that is274302450bytes; the pin'sUPSQUAD_EMBED_ARTIFACT_MODEL_BYTES=274290656is the weights blob only, and the 11,794-byte difference is the license and params layers. Both are correct answers to different questions; provision against this one. - RAM resident is
/api/pssize — what the runner actually mapped, including the KV cache and the compute graph. It is larger than the file on disk for both models, which is the number that matters when sizing a box and the one an operator most often gets wrong by reading the download size. - Bulk ingest is one request carrying all 40 documents, three rounds, median of the per-round rates. This is the ingest leg.
- Query path is 20 individual requests, one short query each, nearest-rank percentiles. This is the embedding leg of a search only — it excludes ANN search, BM25, RRF and recency fusion. It is the part that changes when the model changes, which is the part a model choice is accountable for.
Warm-up is excluded from every figure. The first call pays the model load, which an operator pays once at boot and not per embedding; folding it in is the easiest way to publish a wrong steady-state latency.
The quality tier has no canonical artifact, and that changes what can be published
Upstream Alibaba-NLP/gte-Qwen2-1.5B-instruct ships no GGUF. Every GGUF is a
third-party requantization, and the quantization an operator picks moves the disk
figure by 4.7× on its own.
Artifact provenance, stated because silence reads as endorsement. The quality-tier GGUF is community-produced, unvalidated by us, and NOT digest-pinned under ADR-0029 — unlike the shipped default, which is vendored into the image build and pinned to
r=970aa74c0a90. Measuring an artifact is not endorsing it. Required by founder ratification of PRD #996 v2.2.
Registry-reported sizes from
second-state/gte-Qwen2-1.5B-instruct-GGUF, the repo measured above:
| quantization | disk | quantization | disk | |
|---|---|---|---|---|
| Q2_K | 752.4 MB | Q5_K_M | 1284.8 MB | |
| Q3_K_M | 923.9 MB | Q6_K | 1463.4 MB | |
| Q4_K_M | 1116.7 MB ← measured | Q8_0 | 1893.6 MB | |
| Q4_K_S | 1071.0 MB | f16 | 3558.6 MB |
So there is no single true footprint for this row, and a table that printed one would be asserting a fact about an artifact the operator has not chosen yet. The measured row above is Q4_K_M and says so. RAM and latency were measured on that artifact and do not transfer to another quantization; disk for the others is the registry's own byte count, which is a measurement of the file and not of a run.
FR-19's "~10×" — settled in PRD v2.2, not an open discrepancy
This file previously recorded the "~10×" figure as a discrepancy against a
founder-approved PRD. It is resolved. Founder ratification
5521553923
(2026-09-03T06:30:40Z) amended FR-19 to name the basis rather than renumber the
figure: "~10×" was correct and conservative at equal precision; FR-19 was
under-specified, not wrong. Quote a ratio only with its basis attached:
| basis | ratio | what it is for |
|---|---|---|
| f16 → f16 (like-for-like) | 12.97× | the honest comparison. Both sides at the same precision — the shipped default is f16 (274302450 B / 137M params = 2.002 bytes/param), and the parameter ratio 1.78B / 137M = 12.99× corroborates it independently of file size |
| Q4_K_M (deployable) | 4.07× | what an operator actually installs — and a mixed-precision comparison (f16 default against a Q4 quantization), which is why it is not the like-for-like number |
| query path | 5.31× | the number that actually decides the default |
12.97× is 3558.6 MB ÷ 274.3 MB = 12.9733; the "13.0× on disk" quoted elsewhere in this file is the same measurement at coarser rounding.
The requirement's conclusion is unaffected and in fact strengthened — the real cost is on the synchronous query path, 5.31× at p50, recurring on every search forever, where disk is paid once.
Withdrawn on ratification: the rationale "multi-GB weights that raise the on-prem hardware minimum." Measured at deployable quantization the quality tier is 1116.7 MB disk / 1566.7 MB resident, which raises nobody's hardware floor. It was withdrawn outright rather than softened, because as written it would talk on-prem operators out of an option they can trivially afford.
What was not measured, and why it is not an estimate
- Retrieval quality of the quality tier. Nothing here measures quality; that
is the recall benchmark, and the two must not be
conflated — a model can be fast and blind.
gte-Qwen2uses last-token pooling and it is not established that this GGUF path applies it correctly, so a quality number taken here would be unsound. No quality claim is made for the quality tier anywhere in this milestone. - The hosted 1536-d arm. Not measured — no provider credential and no egress route on the measurement host. #2795 is closed as WAIVED on founder ratification (2026-09-03T06:30:40Z), not pending: the cost declined was a new vendor relationship and a new secret surface, not the ~60 calls. The bound accepted in exchange is permanent — no local-vs-hosted comparative claim, anywhere. See the recall benchmark's binding 4 for the trigger that would reopen the measurement.
Reproducing
# the shipped default
scripts/embedding-footprint-bench.py \
--model nomic-embed-text:v1.5 --label local-nomic \
--base-url http://127.0.0.1:11434 \
--out docs/measurements/embedding-footprint/footprint-local-nomic.json
# the quality tier, at the quantization you intend to deploy
scripts/embedding-footprint-bench.py \
--model 'hf.co/second-state/gte-Qwen2-1.5B-instruct-GGUF:Q4_K_M' \
--label quality-tier-gte-qwen2-q4km \
--base-url http://127.0.0.1:11434 \
--out docs/measurements/embedding-footprint/footprint-quality-tier-gte-qwen2.json
--base-url is the server root, not the /v1 path: disk and RAM come from
Ollama's /api/tags and /api/ps, which have no OpenAI-compatible equivalent.
Against a non-Ollama endpoint the script records both as null with a stated
reason rather than guessing — NFR-3 forbids estimates, and an honest hole is
worth more than a plausible number.
These numbers have to be retaken when the pin moves. A digest bump changes
r=, which makes every existing vector foreign and forces a full re-embed
(ADR-0031 D5.4) — the model has changed, so its footprint is unmeasured again.
That is the same reason the recall benchmark is re-runnable.