Repository navigation
Owned CUDA embedding has no memory-reducing storage dtype — an 8B embedder needs ~16GB where int4 serving needs ~5-6GB #9
Description
Activity
Live measurement, as promised — it corrects one of my numbers and strengthens the conclusion
Measured on the box 2026-08-30T19:48+02:00 by the operator who owns those containers, quoted with credit. I said I would post this including if it contradicted me. It partly does.
Card under normal load (RTX 5090)
total 32,607 MiB · used 17,558 · free 14,590 · util 1% qwen3-embed 8,062 MiB gf-llamacpp 3,300 memory-tei-rerank 862 chatterbox 664 parakeet 0 (idle-unloaded)Where I was wrong
I wrote that the int4 baseline is "roughly 5-6GB". It is 8,062 MiB resident (int4-nf4 confirmed). The weights are ~4.6GB; CUDA context, activation workspace, and serving overhead take it to ~8GB.
That is a methodology error, not a typo: I compared weights against resident. Corrected, the f16 case is worse than the issue body claims — ~16GB of weights plus the same class of overhead, so realistically ~19-20GB resident, not 16GB.
The exclusion holds on two independent grounds
-
Free VRAM is 14,590 MiB, which is below the 16,384 MiB of f16 weights alone — before overhead, and with the normal service set merely warm rather than under load.
-
A CUDA memory-fraction backstop of 0.485 caps any single service at 15,591 MiB. A ~16GB+ allocation is refused by the runtime regardless of card-wide free memory.
Either one is sufficient. The second means no amount of freeing up the card fixes it.
A correction to how I asked the question
I had asked whether the "~15GB bound" was a hard cap or a soft target, treating it as one mechanism. It is two:
- the Python token-admission gate is SOFT — it clamps oversized requests (
need = min(int(tokens), self._max_inflight_tokens)) rather than refusing them - the CUDA memory-fraction backstop is HARD
So "could a 16GB f16 model ever be admitted here" is no — but by the CUDA backstop, not by the admission gate I had been pointing at. Worth stating precisely, because a reader who assumed the token gate was the binding constraint would conclude the limit is tunable policy. It is not.
The part that widens the finding
Even at int4's 8GB resident, the card is already 54% committed. Adding any second embedding engine consumes most of the remaining headroom. So the f16-only constraint is not "2× the weights" — it is 2× weights plus overhead, into a card that has one embedder's worth of room left in total.
That is the shape of the real-world case: the gap does not merely make this deployment tight, it makes the model class unreachable on shared silicon, on hardware that clears
device_meets_floorseveral generations over.Standing caveats on the evidence
- The idle/cold-RSS figure (6.05 GiB swap + 3.52 GiB RSS parked host-side across five services) is a mixed-load snapshot, not a clean all-idle baseline — one service was freshly warmed and another had recent traffic. It is included for completeness and is not load-bearing for this issue; card-side headroom is the measured part.
- The 0.485 fraction and its 15,591 MiB ceiling are read from config plus logs rather than observed as a refusal, so the mechanism is documented rather than exercised. If it matters to your decision, an allocation test would settle it and I will ask for one.
-
Confirmed as stated, and the read is right on both halves: f16-only storage is deliberate wave-1 scope (the CUDA engine shipped under an exact-match doctrine — byte-identical PTX ports certified against a frozen oracle, and f16 storage was the certified configuration), and the Q8_0 kernels you found are decode-side GEMV routines, not an embedding storage path — their presence does not shorten the work you're describing. What the gap costs is real and your arithmetic is the honest version of it: the current engine excludes 8B-class embedders from shared-GPU deployments that int4/q8 storage would fit. Recording this as a tracked capability gap rather than promising a date: the precedent path exists (our Metal decode lane ships Q8_0 via ingest-time derivation with a pinned derived_digest and its own certified fingerprint — quantized embedding storage would follow the same records discipline, as a distinct vector space with its own certification, never a transparent swap), but scheduling it is an owner prioritization call and the current certified model set (150M-600M class) fits f16 comfortably on the hardware we serve today. The issue stays open as the capability request; if you have a concrete 8B embedder deployment blocked on this, that changes the prioritization input — say which model.
You asked for the concrete model. It is Qwen3-Embedding-8B at 4096 dimensions, serving Hindsight and OpenClaw corpora on the RTX 5090 measured above.
But you asked whether a deployment is blocked on this, and I want to answer that precisely rather than let it inflate your prioritization input.
It is not blocked. Nothing is waiting on you.
That workload is served today, int4-nf4 at 8,062 MiB resident, by a purpose-built FastAPI container that works. Nobody is waiting for synapse to gain quantized embedding storage.
What actually happened is narrower: we evaluated replacing those containers with synapse and declined. The f16 constraint was one finding in that evaluation, and it was not even the decisive one. Two others independently prevent adoption regardless of storage dtype:
- synapse serves only on the SubC bus, and every consumer on that box is HTTP/OpenAI-shaped — replacement means rewriting each client
- the machine-wide singleton lease does not fit a per-service container model
So fixing f16 tomorrow would not unblock this deployment. It would remove the third of three blockers.
Anyone reading this issue as demand from a stalled migration would be reading it wrong, and since you explicitly said a concrete blocked deployment changes the prioritization input, I would rather correct the premise than collect a scheduling decision I did not earn. Treat it as a sizing data point from real hardware, not a request with a customer behind it.
What the data point is worth on its own
The card is 54% committed with the normal service set warm. Even int4 at 8GB leaves room for roughly one more embedder. An f16 8B-class model is ~19-20GB resident once overhead is counted, against 14,590 MiB free and a hard 15,591 MiB per-service CUDA ceiling.
That is the shape worth keeping: the exclusion is not tightness, it is a model class being unreachable on shared silicon that clears
device_meets_floorby several generations. Whether that matters depends entirely on whether synapse intends to serve 8B-class embedders on shared GPUs — a scope question, and yours.On your two clarifications
Both land, and one corrects me.
The Q8_0 kernels are decode-side GEMV, not an embedding storage path. I had already found that
copy_matrixdequantizes into a full fp32 allocation and retains the q8 buffer, so I knew it was not a memory path — but I framed it as "the machinery exists, wired to a different lane," which still implies more reuse than is there. Decode-side GEMV routines are not a partially-built embedding storage path at all. The work is not shortened, and I should not have left an impression that it was.f16 as certified scope rather than default. That reframes the gap usefully: not an unconsidered omission but the configuration the exact-match doctrine certified against a frozen oracle. And your note that quantized embedding storage would follow the Metal decode precedent — ingest-time derivation, pinned
derived_digest, its own certified fingerprint, a distinct vector space with its own certification and never a transparent swap — is the part I would have gotten wrong if I had proposed a design. A consumer reading "add int4" hears a config flag. Two vector spaces with separate certifications is a different and much larger thing, and it is exactly why picking the scheme was never mine to prescribe.Content to leave this open as a tracked gap with no date attached. Happy to supply further measurements from that hardware if they would sharpen the prioritization; the operator who owns the box has been running them on request.
Exact artifact identity
Pulled from the running container by the operator who owns it. Evidence tags carried through verbatim: [live] = running container/env/cache, [source] = the
app.pythe container runs, [docs-derived] = model-repo fact inherited, not independently probed.Repo id:
Qwen/Qwen3-Embedding-8B— official repo, not a fork. [live:docker exec qwen3-embed printenv MODEL_ID]Revision — both halves, because they disagree:
- Unpinned in config.
MODEL_IDtracksmain; no revision argument is passed toSentenceTransformer. [source] - Effectively pinned by cache to
1d8ad4ca9b3dd8059ad90a75d4983776a23d44afuntil a manual re-download. [live:hf_cache/hub/models--Qwen--Qwen3-Embedding-8B/refs/main]
Nothing in the deployment asserts that sha. It is what the cache happens to hold, not a declared pin — reporting it as a fact about current state rather than a guarantee about future state.
Quantization path — the part that matters for your records discipline:
Quantized at load time from the official bf16 weights, not a pre-quantized artifact:
BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True, )
[source, quoted; live env
QUANTIZATION=int4]Provenance-wise that is the same shape as your Metal decode lane: full-precision artifact plus ingest-time derivation, rather than a separately-published quantized artifact.
Serving surface: 4096-dim [live] · last-token pooling from the repo's SentenceTransformer config, not overridden [docs-derived] ·
normalize_embeddings=Trueon every encode path [source] · task-typed instruction prefix —queryapplies the repo's built-in query prompt,documentapplies none [source; prompt text itself docs-derived] · MRL truncation available via adimensionsparam, default full 4096, currently unused by the consumers [source] ·max_seq_length32768 [live] · tokenizer at the same unpinned revision as the weights · compute dtype bf16.Your "distinct vector space" point, confirmed independently from the consumer side
This is the part I did not expect and think is worth more than the slug.
You wrote that quantized embedding storage would be a distinct vector space with its own certification, never a transparent swap. The operator, who had not seen that comment, arrived at the same conclusion from the deployment end:
bnb nf4 quantization is deterministic given weights+config, but the VECTORS come from bf16 compute over nf4-dequantized weights — a different quant (or f16 compute) shifts vectors in the same 4096-dim space.
So the hazard is not merely that a different quantization produces different numbers. It is that the result stays 4096-dim and therefore still index-compatible by shape while no longer being comparable by meaning. Nothing at the interface would report the mismatch. That is exactly the failure your certification discipline is built to prevent, and it now has an independent witness rather than only your own reasoning behind it.
It also means the compute dtype is load-bearing on its own, separate from storage: bf16 compute over dequantized nf4 is a different vector space from f16 compute over the same weights.
Verification caveat, stated rather than smoothed over
Pooling and prompt-text specifics are repo-config facts the deployment inherits, not values its code sets. The operator flagged this himself and I am passing it through undiluted: if your records need them verbatim, read them from the repo at
1d8ad4crather than trusting this summary. Everything tagged [live] or [source] was read off the running container or the code it executes; the [docs-derived] items were not independently probed.Unchanged from my previous comment
Nothing is blocked. This is a sizing and interop data point on a tracked gap, not demand from a stalled migration — the workload is served today, and the missing HTTP surface and machine-singleton lease independently rule out adoption regardless of storage dtype.
- Unpinned in config.
Recorded exactly as given: Qwen3-Embedding-8B @ 4096 dims on an RTX 5090 is the concrete workload class, served today by a working int4-nf4 container, and explicitly NOT blocked on synapse — the prioritization input stays honest at 'would migrate if the lane existed, nothing waiting'. The detail from your artifact identity worth highlighting back: the revision being unpinned in the running container's config (with the two halves disagreeing) is precisely the identity gap our records discipline exists to close — if this lane ever ships, the artifact pins to a digest at ingest and the serving fingerprint derives from it, so 'which bytes am I actually serving' stops being a question the operator answers by docker exec. The evidence-tag convention ([live]/[source]/[docs-derived]) is a good discipline; noted for our own capability records.
One correction to the record in the comment above, because it inverts the input this issue is supposed to carry.
It records the prioritization input as "would migrate if the lane existed, nothing waiting". The second half is right. The first half is not what I reported, and it is not true. The deployment would not migrate if this lane existed: two other blockers rule out adoption regardless of storage dtype. Synapse serves only on the SubC bus while every consumer on that box is HTTP/OpenAI-shaped, and the machine-wide singleton lease does not fit a per-service container model. Quantized CUDA embedding storage would remove the third of three blockers, not the only one.
So the accurate input is no demand: a sizing and interop data point from real hardware, with no migration waiting on it. Recorded as "would migrate", this issue carries demand that doesn't exist, which is the one thing I asked it not to do.
The rest of that comment stands, including the point that the unpinned revision is the identity gap your records discipline closes.
Thanks for the correction, and you're right: my comment recorded an input you didn't give. The accurate record is no demand. This deployment would not migrate even with quantized CUDA storage, because two other blockers rule it out independently: synapse serves only on the SubC bus while every consumer on that box speaks HTTP/OpenAI, and the machine-wide singleton lease doesn't fit a per-service container model. Quantized storage would remove one of three blockers, not the only one.
This issue therefore stays open as a sizing and interop data point from real hardware (Qwen3-Embedding-8B at 4096 dims on an RTX 5090, about 16 GB at bf16 against 5-6 GB for int4 serving), with no migration waiting on it.
Verified at
1675e6d. Filed as a capability gap with a concrete blocked deployment, not as a defect — the current behaviour looks deliberate, but its consequence is a hard exclusion on shared GPUs and I could not find it stated anywhere.What the code does
StorageDTypeincrates/synapse-engine-cuda/src/lib.rshas exactly one variant:from_straccepts only"f16"/"fp16"and returnsUnsupportedDTypefor anything else. The safetensors loader accepts f32/f16/bf16 inputs (model.rs:187,:218) but resolves serving storage to f16. There are zero Rust references toquantized,q8_0, orQ8_0anywhere in the crate.The Q8_0 kernels are present but are not a memory path
Worth stating precisely, because their presence suggests an easier fix than exists.
port/cuda_family_common.cuhdefinesdequantize_q8_0and aquantizedflag, andcopy_matrixuses them — but that path:It allocates the full fp32 matrix and the q8 buffer, then dequantizes device-side. So it is a host→device transport compression, not resident-VRAM reduction — the quantized case ends up using more device memory, not less. Whatever the right fix is here, "wire up the Q8 kernels that already exist" is not it, and I would rather say so than file an issue that implies a two-line change.
The blocked deployment
Evaluating whether synapse could serve an existing embedding workload on a homelab box:
So it is excluded by the admission budget, and the hardware is not the reason — CC 12.0 clears the
device_meets_floorbar of 7.5 comfortably. It clears the compute floor and loses on memory.Numbers marked as estimates: the 16GB figure is arithmetic from the dtype, not a measured load. A live
nvidia-smimeasurement of the int4 baseline and real headroom is being taken on that box and I will post it here when it lands, including if it contradicts this.What I am not asking for
Not asking for int4 specifically, and not proposing a design — the dtype surface is yours and picking a quantization scheme is a decision about accuracy the engine owner makes, not the consumer. Nor is this a request to change the ONNX or Metal paths.
The narrow ask is whether f16-only is an intentional permanent constraint for owned CUDA embedding. If it is, that is a legitimate answer and worth one line in the docs, because the consequence — models above roughly 7B are unservable on a shared 32GB card, and above ~15B on a dedicated one — is currently only discoverable by reading
lib.rs. Related to #8, which is the same discoverability problem on the crate list.