Skip to content

Add clear_kv_cache to quantized gemma3, phi3, glm4 and lfm2 - #3709

Open
kitecosmic wants to merge 1 commit into
huggingface:mainfrom
kitecosmic:add-clear-kv-cache-quantized-models
Open

kitecosmic wants to merge 1 commit into
huggingface:mainfrom
kitecosmic:add-clear-kv-cache-quantized-models

Conversation

@kitecosmic

Copy link
Copy Markdown

Adds a public clear_kv_cache() to the ModelWeights of quantized_gemma3, quantized_phi3, quantized_glm4 and quantized_lfm2, mirroring the method already exposed by quantized_llama, quantized_qwen2 and quantized_qwen3.

Why

When a single loaded quantized model serves multiple independent requests, the caller needs an explicit way to drop cached attention state between conversations:

  • gemma3/phi3/glm4 currently reset their cache only implicitly, when forward is called with index_pos == 0. That works, but it is an undocumented contract — downstream code that must guarantee request isolation (no tokens leaking across calls) has nothing explicit to call, unlike with the llama/qwen2/qwen3 siblings.
  • quantized_lfm2 is the strongest case: ShortConvLayer::forward ignores index_pos entirely, so its conv state has no implicit reset at all — a new conversation whose first forward has seq_len == 1 silently reuses the previous conversation's conv state. Its clear_kv_cache clears both the attention KV pairs and the conv state.
  • Parity across the quantized model family lets downstream code hold one enum over ModelWeights variants and clear state uniformly.
  • For quantized_gemma3 (per-layer Option<(Tensor, Tensor)>), the explicit clear also frees the cached tensors between requests instead of keeping the last conversation's KV alive until the next forward.

The implementation copies the exact pattern of quantized_llama::ModelWeights::clear_kv_cache (iterate layers, clear the K/V pair): kv_cache = None for gemma3, KvCache::reset() for phi3/glm4 (which use candle_nn::kv_cache::KvCache).

No behavior change for existing callers — the method is additive. Same pattern previously merged for the siblings in #3536 (quantized llama/qwen2) and #3189 (quantized qwen3).

Related CPU findings from the same embedding work, filed separately: #3707 (no batch path in the quantized matmul — prefill runs at decode speed) and #3708 (no lazy/mmap loading for GGUF).

kitecosmic added a commit to kitecosmic/synsema that referenced this pull request Sep 22, 2026
…d without candle, and architectures declared in a file instead of a PR

candle belongs to Hugging Face, which NVIDIA agreed to buy on 2026-09-02, and the dependency
already hurt without anyone acting in bad faith: a three-line PR of ours (huggingface/candle#3709)
has been approved and idle since July, and with it `quantized_gemma3` and `quantized_qwen3_moe`
stay blocked for our pool. The pattern that hits us is not "candle lacks X", it is "candle has X
in full and the quantized version arrives late or never" — and we live entirely in quantized.

So the calculation moved out of `llm_local.rs` into `engine/crates/synsema-infer`, a crate with
three doors (`generate`, `embed`, `decide`) and candle behind a trait of ours. `llm_local.rs` is
now the adapter between the LLM protocol and that crate, and no other crate in the engine names a
candle type. The full reasoning, the tandas and the closing criteria are in
`specs/synsema-infer.md`.

What a user gets:

- `SYNSEMA_LLM_MODEL` takes a name, not only a path: a `.gguf` path, a `model:tag` already in the
  Ollama cache, or an `org/repo` already in the Hugging Face cache. Nothing is ever downloaded.
  Ollama's store is content-addressed, so the sha256 of the weights comes for free and is
  reported as provenance.

- A backend written by us, `SYNSEMA_INFER_BACKEND=rust`, with no candle in its tree. It runs
  gemma3 — which candle does not ship quantized — and it picks AVX/AVX2+FMA/AVX-512/NEON at run
  time, so the official binary uses the instructions of the machine it lands on. candle stays the
  default while both exist: switching engines changes the generated text, which makes it part of
  what you declare to reproduce an output.

- Architectures are a file. `.archdef` is a flat list of named steps over the tensors of a GGUF;
  the four we ship (llama, qwen2, qwen3, gemma3) are embedded, and `SYNSEMA_INFER_ARCHDEF` points
  at a directory that adds new ones or replaces ours **without recompiling anything**. The four
  handwritten architectures were migrated to the format and produce the same bits — prompt,
  incremental decode and after clearing the KV cache. The grammar has no conditionals, no loops
  and no way to touch a file, a socket or the environment, so running someone else's definition
  does not run their code; a file that fails to parse leaves that architecture unavailable with
  the error of the file, and never falls back to ours in silence.

- `judge` runs locally: `SYNSEMA_JUDGE_PROVIDER=laya` answers `whether`, `choose` and `rate`
  against a Laya (ModernBERT) checkpoint on disk, with no network, no secret and no cost per
  token. The official binaries now ship with it compiled in — a capability the docs describe and
  the released binary lacks is not a capability.

- `synsema llm status` says what is on this machine: models already downloaded with their origin
  and sha, the architectures **of the backend that will actually run** with their origin and sha,
  and the definitions that failed to load with their error. `--json` carries the same under
  `inference`, with the full sha.

Three bugs that the declarative path exposed, all of them previously taken as correct:

1. The SentencePiece tokenizer was wrong for the whole llama family. Those GGUFs store ranks, not
   log-probabilities, and we segmented them with a Viterbi pass maximising the sum of scores, so
   `The capital of France is` entered the model as eleven fragments instead of five words. It is
   now the reference algorithm, written by us. The BPE family (qwen, llama 3) was never affected.
2. gemma3's MLP activation was `silu` and it is `gelu_pytorch_tanh` — copied from candle, whose
   quantized gemma3 hardcodes it while its own non-quantized one reads the config.
3. Gemma had no chat template, so it fell back to plain mode and behaved like a base model.

The lesson is the reason they survived: an oracle proves two implementations agree, not that they
are right. I4 validated gemma3 against candle and it was green; candle has the same error in two of
the three cases. The engine now generates token for token what Ollama generates for the same
prompt, and where we can we contrast against llama.cpp and not only against candle.

Tests: 365 green across the three crates touched (synsema-infer 213, synsema-runtime 125,
synsema-cli 27), plus a live probe with real weights for every claim above.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant