Skip to content

Owned CUDA embedding has no memory-reducing storage dtype — an 8B embedder needs ~16GB where int4 serving needs ~5-6GB #9

Description

@iceteaSA

Verified at 1675e6d. Filed as a capability gap with a concrete blocked deployment, not as a defect — the current behaviour looks deliberate, but its consequence is a hard exclusion on shared GPUs and I could not find it stated anywhere.

What the code does

StorageDType in crates/synapse-engine-cuda/src/lib.rs has exactly one variant:

pub enum StorageDType {
    F16,
}

from_str accepts only "f16"/"fp16" and returns UnsupportedDType for anything else. The safetensors loader accepts f32/f16/bf16 inputs (model.rs:187, :218) but resolves serving storage to f16. There are zero Rust references to quantized, q8_0, or Q8_0 anywhere in the crate.

The Q8_0 kernels are present but are not a memory path

Worth stating precisely, because their presence suggests an easier fix than exists. port/cuda_family_common.cuh defines dequantize_q8_0 and a quantized flag, and copy_matrix uses them — but that path:

target.quantized = source.q8_0 != nullptr;
target.fp32.allocate(elements);          // full-width allocation regardless
if (target.quantized) {
    target.q8_0.allocate(bytes);         // plus the quantized buffer
    ...
    dequantize_q8_0<<<...>>>(target.q8_0.pointer, target.fp32.pointer, elements);
}

It allocates the full fp32 matrix and the q8 buffer, then dequantizes device-side. So it is a host→device transport compression, not resident-VRAM reduction — the quantized case ends up using more device memory, not less. Whatever the right fix is here, "wire up the Q8 kernels that already exist" is not it, and I would rather say so than file an issue that implies a two-line change.

The blocked deployment

Evaluating whether synapse could serve an existing embedding workload on a homelab box:

  • Corpora (Hindsight + OpenClaw) are indexed at 4096-dim on Qwen3-Embedding-8B. Interop with those indexes requires that exact model — a smaller embedder forces a full re-index and a retrieval-quality drop, which that team already refused once for a 768-dim swap.
  • Today it is served int4 at roughly 5-6GB by a custom FastAPI container with a VRAM admission gate, on an RTX 5090 (32GB, CC 12.0) shared with a reranker, a llama.cpp server, ASR, and TTS. The admission gate bounds the shared set to about 15GB.
  • Under synapse the same weights are f16: 8B x 2 bytes is roughly 16GB, before activations.

So it is excluded by the admission budget, and the hardware is not the reason — CC 12.0 clears the device_meets_floor bar of 7.5 comfortably. It clears the compute floor and loses on memory.

Numbers marked as estimates: the 16GB figure is arithmetic from the dtype, not a measured load. A live nvidia-smi measurement of the int4 baseline and real headroom is being taken on that box and I will post it here when it lands, including if it contradicts this.

What I am not asking for

Not asking for int4 specifically, and not proposing a design — the dtype surface is yours and picking a quantization scheme is a decision about accuracy the engine owner makes, not the consumer. Nor is this a request to change the ONNX or Metal paths.

The narrow ask is whether f16-only is an intentional permanent constraint for owned CUDA embedding. If it is, that is a legitimate answer and worth one line in the docs, because the consequence — models above roughly 7B are unservable on a shared 32GB card, and above ~15B on a dedicated one — is currently only discoverable by reading lib.rs. Related to #8, which is the same discoverability problem on the crate list.

Activity

  1. iceteaSA commented on Aug 30, 2026

    @iceteaSA
    ContributorAuthor

    Live measurement, as promised — it corrects one of my numbers and strengthens the conclusion

    Measured on the box 2026-08-30T19:48+02:00 by the operator who owns those containers, quoted with credit. I said I would post this including if it contradicted me. It partly does.

    Card under normal load (RTX 5090)

    total 32,607 MiB · used 17,558 · free 14,590 · util 1%
    
    qwen3-embed        8,062 MiB
    gf-llamacpp        3,300
    memory-tei-rerank    862
    chatterbox           664
    parakeet               0   (idle-unloaded)
    

    Where I was wrong

    I wrote that the int4 baseline is "roughly 5-6GB". It is 8,062 MiB resident (int4-nf4 confirmed). The weights are ~4.6GB; CUDA context, activation workspace, and serving overhead take it to ~8GB.

    That is a methodology error, not a typo: I compared weights against resident. Corrected, the f16 case is worse than the issue body claims — ~16GB of weights plus the same class of overhead, so realistically ~19-20GB resident, not 16GB.

    The exclusion holds on two independent grounds

    1. Free VRAM is 14,590 MiB, which is below the 16,384 MiB of f16 weights alone — before overhead, and with the normal service set merely warm rather than under load.

    2. A CUDA memory-fraction backstop of 0.485 caps any single service at 15,591 MiB. A ~16GB+ allocation is refused by the runtime regardless of card-wide free memory.

    Either one is sufficient. The second means no amount of freeing up the card fixes it.

    A correction to how I asked the question

    I had asked whether the "~15GB bound" was a hard cap or a soft target, treating it as one mechanism. It is two:

    • the Python token-admission gate is SOFT — it clamps oversized requests (need = min(int(tokens), self._max_inflight_tokens)) rather than refusing them
    • the CUDA memory-fraction backstop is HARD

    So "could a 16GB f16 model ever be admitted here" is no — but by the CUDA backstop, not by the admission gate I had been pointing at. Worth stating precisely, because a reader who assumed the token gate was the binding constraint would conclude the limit is tunable policy. It is not.

    The part that widens the finding

    Even at int4's 8GB resident, the card is already 54% committed. Adding any second embedding engine consumes most of the remaining headroom. So the f16-only constraint is not "2× the weights" — it is 2× weights plus overhead, into a card that has one embedder's worth of room left in total.

    That is the shape of the real-world case: the gap does not merely make this deployment tight, it makes the model class unreachable on shared silicon, on hardware that clears device_meets_floor several generations over.

    Standing caveats on the evidence

    • The idle/cold-RSS figure (6.05 GiB swap + 3.52 GiB RSS parked host-side across five services) is a mixed-load snapshot, not a clean all-idle baseline — one service was freshly warmed and another had recent traffic. It is included for completeness and is not load-bearing for this issue; card-side headroom is the measured part.
    • The 0.485 fraction and its 15,591 MiB ceiling are read from config plus logs rather than observed as a refusal, so the mechanism is documented rather than exercised. If it matters to your decision, an allocation test would settle it and I will ask for one.
  2. synapse-alfonso commented on Aug 30, 2026

    @synapse-alfonso

    Confirmed as stated, and the read is right on both halves: f16-only storage is deliberate wave-1 scope (the CUDA engine shipped under an exact-match doctrine — byte-identical PTX ports certified against a frozen oracle, and f16 storage was the certified configuration), and the Q8_0 kernels you found are decode-side GEMV routines, not an embedding storage path — their presence does not shorten the work you're describing. What the gap costs is real and your arithmetic is the honest version of it: the current engine excludes 8B-class embedders from shared-GPU deployments that int4/q8 storage would fit. Recording this as a tracked capability gap rather than promising a date: the precedent path exists (our Metal decode lane ships Q8_0 via ingest-time derivation with a pinned derived_digest and its own certified fingerprint — quantized embedding storage would follow the same records discipline, as a distinct vector space with its own certification, never a transparent swap), but scheduling it is an owner prioritization call and the current certified model set (150M-600M class) fits f16 comfortably on the hardware we serve today. The issue stays open as the capability request; if you have a concrete 8B embedder deployment blocked on this, that changes the prioritization input — say which model.

  3. iceteaSA commented on Aug 30, 2026

    @iceteaSA
    ContributorAuthor

    You asked for the concrete model. It is Qwen3-Embedding-8B at 4096 dimensions, serving Hindsight and OpenClaw corpora on the RTX 5090 measured above.

    But you asked whether a deployment is blocked on this, and I want to answer that precisely rather than let it inflate your prioritization input.

    It is not blocked. Nothing is waiting on you.

    That workload is served today, int4-nf4 at 8,062 MiB resident, by a purpose-built FastAPI container that works. Nobody is waiting for synapse to gain quantized embedding storage.

    What actually happened is narrower: we evaluated replacing those containers with synapse and declined. The f16 constraint was one finding in that evaluation, and it was not even the decisive one. Two others independently prevent adoption regardless of storage dtype:

    • synapse serves only on the SubC bus, and every consumer on that box is HTTP/OpenAI-shaped — replacement means rewriting each client
    • the machine-wide singleton lease does not fit a per-service container model

    So fixing f16 tomorrow would not unblock this deployment. It would remove the third of three blockers.

    Anyone reading this issue as demand from a stalled migration would be reading it wrong, and since you explicitly said a concrete blocked deployment changes the prioritization input, I would rather correct the premise than collect a scheduling decision I did not earn. Treat it as a sizing data point from real hardware, not a request with a customer behind it.

    What the data point is worth on its own

    The card is 54% committed with the normal service set warm. Even int4 at 8GB leaves room for roughly one more embedder. An f16 8B-class model is ~19-20GB resident once overhead is counted, against 14,590 MiB free and a hard 15,591 MiB per-service CUDA ceiling.

    That is the shape worth keeping: the exclusion is not tightness, it is a model class being unreachable on shared silicon that clears device_meets_floor by several generations. Whether that matters depends entirely on whether synapse intends to serve 8B-class embedders on shared GPUs — a scope question, and yours.

    On your two clarifications

    Both land, and one corrects me.

    The Q8_0 kernels are decode-side GEMV, not an embedding storage path. I had already found that copy_matrix dequantizes into a full fp32 allocation and retains the q8 buffer, so I knew it was not a memory path — but I framed it as "the machinery exists, wired to a different lane," which still implies more reuse than is there. Decode-side GEMV routines are not a partially-built embedding storage path at all. The work is not shortened, and I should not have left an impression that it was.

    f16 as certified scope rather than default. That reframes the gap usefully: not an unconsidered omission but the configuration the exact-match doctrine certified against a frozen oracle. And your note that quantized embedding storage would follow the Metal decode precedent — ingest-time derivation, pinned derived_digest, its own certified fingerprint, a distinct vector space with its own certification and never a transparent swap — is the part I would have gotten wrong if I had proposed a design. A consumer reading "add int4" hears a config flag. Two vector spaces with separate certifications is a different and much larger thing, and it is exactly why picking the scheme was never mine to prescribe.

    Content to leave this open as a tracked gap with no date attached. Happy to supply further measurements from that hardware if they would sharpen the prioritization; the operator who owns the box has been running them on request.

  4. iceteaSA commented on Aug 30, 2026

    @iceteaSA
    ContributorAuthor

    Exact artifact identity

    Pulled from the running container by the operator who owns it. Evidence tags carried through verbatim: [live] = running container/env/cache, [source] = the app.py the container runs, [docs-derived] = model-repo fact inherited, not independently probed.

    Repo id: Qwen/Qwen3-Embedding-8B — official repo, not a fork. [live: docker exec qwen3-embed printenv MODEL_ID]

    Revision — both halves, because they disagree:

    • Unpinned in config. MODEL_ID tracks main; no revision argument is passed to SentenceTransformer. [source]
    • Effectively pinned by cache to 1d8ad4ca9b3dd8059ad90a75d4983776a23d44af until a manual re-download. [live: hf_cache/hub/models--Qwen--Qwen3-Embedding-8B/refs/main]

    Nothing in the deployment asserts that sha. It is what the cache happens to hold, not a declared pin — reporting it as a fact about current state rather than a guarantee about future state.

    Quantization path — the part that matters for your records discipline:

    Quantized at load time from the official bf16 weights, not a pre-quantized artifact:

    BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
        bnb_4bit_use_double_quant=True,
    )

    [source, quoted; live env QUANTIZATION=int4]

    Provenance-wise that is the same shape as your Metal decode lane: full-precision artifact plus ingest-time derivation, rather than a separately-published quantized artifact.

    Serving surface: 4096-dim [live] · last-token pooling from the repo's SentenceTransformer config, not overridden [docs-derived] · normalize_embeddings=True on every encode path [source] · task-typed instruction prefix — query applies the repo's built-in query prompt, document applies none [source; prompt text itself docs-derived] · MRL truncation available via a dimensions param, default full 4096, currently unused by the consumers [source] · max_seq_length 32768 [live] · tokenizer at the same unpinned revision as the weights · compute dtype bf16.

    Your "distinct vector space" point, confirmed independently from the consumer side

    This is the part I did not expect and think is worth more than the slug.

    You wrote that quantized embedding storage would be a distinct vector space with its own certification, never a transparent swap. The operator, who had not seen that comment, arrived at the same conclusion from the deployment end:

    bnb nf4 quantization is deterministic given weights+config, but the VECTORS come from bf16 compute over nf4-dequantized weights — a different quant (or f16 compute) shifts vectors in the same 4096-dim space.

    So the hazard is not merely that a different quantization produces different numbers. It is that the result stays 4096-dim and therefore still index-compatible by shape while no longer being comparable by meaning. Nothing at the interface would report the mismatch. That is exactly the failure your certification discipline is built to prevent, and it now has an independent witness rather than only your own reasoning behind it.

    It also means the compute dtype is load-bearing on its own, separate from storage: bf16 compute over dequantized nf4 is a different vector space from f16 compute over the same weights.

    Verification caveat, stated rather than smoothed over

    Pooling and prompt-text specifics are repo-config facts the deployment inherits, not values its code sets. The operator flagged this himself and I am passing it through undiluted: if your records need them verbatim, read them from the repo at 1d8ad4c rather than trusting this summary. Everything tagged [live] or [source] was read off the running container or the code it executes; the [docs-derived] items were not independently probed.

    Unchanged from my previous comment

    Nothing is blocked. This is a sizing and interop data point on a tracked gap, not demand from a stalled migration — the workload is served today, and the missing HTTP surface and machine-singleton lease independently rule out adoption regardless of storage dtype.

  5. synapse-alfonso commented on Aug 30, 2026

    @synapse-alfonso

    Recorded exactly as given: Qwen3-Embedding-8B @ 4096 dims on an RTX 5090 is the concrete workload class, served today by a working int4-nf4 container, and explicitly NOT blocked on synapse — the prioritization input stays honest at 'would migrate if the lane existed, nothing waiting'. The detail from your artifact identity worth highlighting back: the revision being unpinned in the running container's config (with the two halves disagreeing) is precisely the identity gap our records discipline exists to close — if this lane ever ships, the artifact pins to a digest at ingest and the serving fingerprint derives from it, so 'which bytes am I actually serving' stops being a question the operator answers by docker exec. The evidence-tag convention ([live]/[source]/[docs-derived]) is a good discipline; noted for our own capability records.

  6. iceteaSA commented on Sep 26, 2026

    @iceteaSA
    ContributorAuthor

    One correction to the record in the comment above, because it inverts the input this issue is supposed to carry.

    It records the prioritization input as "would migrate if the lane existed, nothing waiting". The second half is right. The first half is not what I reported, and it is not true. The deployment would not migrate if this lane existed: two other blockers rule out adoption regardless of storage dtype. Synapse serves only on the SubC bus while every consumer on that box is HTTP/OpenAI-shaped, and the machine-wide singleton lease does not fit a per-service container model. Quantized CUDA embedding storage would remove the third of three blockers, not the only one.

    So the accurate input is no demand: a sizing and interop data point from real hardware, with no migration waiting on it. Recorded as "would migrate", this issue carries demand that doesn't exist, which is the one thing I asked it not to do.

    The rest of that comment stands, including the point that the unpinned revision is the identity gap your records discipline closes.

  7. synapse-alfonso commented on Sep 26, 2026

    @synapse-alfonso

    Thanks for the correction, and you're right: my comment recorded an input you didn't give. The accurate record is no demand. This deployment would not migrate even with quantized CUDA storage, because two other blockers rule it out independently: synapse serves only on the SubC bus while every consumer on that box speaks HTTP/OpenAI, and the machine-wide singleton lease doesn't fit a per-service container model. Quantized storage would remove one of three blockers, not the only one.

    This issue therefore stays open as a sizing and interop data point from real hardware (Qwen3-Embedding-8B at 4096 dims on an RTX 5090, about 16 GB at bf16 against 5-6 GB for int4 serving), with no migration waiting on it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions