Repository navigation
Add clear_kv_cache to quantized gemma3, phi3, glm4 and lfm2 - #3709
Open
kitecosmic wants to merge 1 commit into
Open
kitecosmic wants to merge 1 commit into
kitecosmic wants to merge 1 commit into
Conversation
kitecosmic
added a commit
to kitecosmic/synsema
that referenced
this pull request
Sep 22, 2026
…d without candle, and architectures declared in a file instead of a PR candle belongs to Hugging Face, which NVIDIA agreed to buy on 2026-09-02, and the dependency already hurt without anyone acting in bad faith: a three-line PR of ours (huggingface/candle#3709) has been approved and idle since July, and with it `quantized_gemma3` and `quantized_qwen3_moe` stay blocked for our pool. The pattern that hits us is not "candle lacks X", it is "candle has X in full and the quantized version arrives late or never" — and we live entirely in quantized. So the calculation moved out of `llm_local.rs` into `engine/crates/synsema-infer`, a crate with three doors (`generate`, `embed`, `decide`) and candle behind a trait of ours. `llm_local.rs` is now the adapter between the LLM protocol and that crate, and no other crate in the engine names a candle type. The full reasoning, the tandas and the closing criteria are in `specs/synsema-infer.md`. What a user gets: - `SYNSEMA_LLM_MODEL` takes a name, not only a path: a `.gguf` path, a `model:tag` already in the Ollama cache, or an `org/repo` already in the Hugging Face cache. Nothing is ever downloaded. Ollama's store is content-addressed, so the sha256 of the weights comes for free and is reported as provenance. - A backend written by us, `SYNSEMA_INFER_BACKEND=rust`, with no candle in its tree. It runs gemma3 — which candle does not ship quantized — and it picks AVX/AVX2+FMA/AVX-512/NEON at run time, so the official binary uses the instructions of the machine it lands on. candle stays the default while both exist: switching engines changes the generated text, which makes it part of what you declare to reproduce an output. - Architectures are a file. `.archdef` is a flat list of named steps over the tensors of a GGUF; the four we ship (llama, qwen2, qwen3, gemma3) are embedded, and `SYNSEMA_INFER_ARCHDEF` points at a directory that adds new ones or replaces ours **without recompiling anything**. The four handwritten architectures were migrated to the format and produce the same bits — prompt, incremental decode and after clearing the KV cache. The grammar has no conditionals, no loops and no way to touch a file, a socket or the environment, so running someone else's definition does not run their code; a file that fails to parse leaves that architecture unavailable with the error of the file, and never falls back to ours in silence. - `judge` runs locally: `SYNSEMA_JUDGE_PROVIDER=laya` answers `whether`, `choose` and `rate` against a Laya (ModernBERT) checkpoint on disk, with no network, no secret and no cost per token. The official binaries now ship with it compiled in — a capability the docs describe and the released binary lacks is not a capability. - `synsema llm status` says what is on this machine: models already downloaded with their origin and sha, the architectures **of the backend that will actually run** with their origin and sha, and the definitions that failed to load with their error. `--json` carries the same under `inference`, with the full sha. Three bugs that the declarative path exposed, all of them previously taken as correct: 1. The SentencePiece tokenizer was wrong for the whole llama family. Those GGUFs store ranks, not log-probabilities, and we segmented them with a Viterbi pass maximising the sum of scores, so `The capital of France is` entered the model as eleven fragments instead of five words. It is now the reference algorithm, written by us. The BPE family (qwen, llama 3) was never affected. 2. gemma3's MLP activation was `silu` and it is `gelu_pytorch_tanh` — copied from candle, whose quantized gemma3 hardcodes it while its own non-quantized one reads the config. 3. Gemma had no chat template, so it fell back to plain mode and behaved like a base model. The lesson is the reason they survived: an oracle proves two implementations agree, not that they are right. I4 validated gemma3 against candle and it was green; candle has the same error in two of the three cases. The engine now generates token for token what Ollama generates for the same prompt, and where we can we contrast against llama.cpp and not only against candle. Tests: 365 green across the three crates touched (synsema-infer 213, synsema-runtime 125, synsema-cli 27), plus a live probe with real weights for every claim above.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a public
clear_kv_cache()to theModelWeightsofquantized_gemma3,quantized_phi3,quantized_glm4andquantized_lfm2, mirroring the method already exposed byquantized_llama,quantized_qwen2andquantized_qwen3.Why
When a single loaded quantized model serves multiple independent requests, the caller needs an explicit way to drop cached attention state between conversations:
forwardis called withindex_pos == 0. That works, but it is an undocumented contract — downstream code that must guarantee request isolation (no tokens leaking across calls) has nothing explicit to call, unlike with the llama/qwen2/qwen3 siblings.quantized_lfm2is the strongest case:ShortConvLayer::forwardignoresindex_posentirely, so its conv state has no implicit reset at all — a new conversation whose first forward hasseq_len == 1silently reuses the previous conversation's conv state. Itsclear_kv_cacheclears both the attention KV pairs and the conv state.ModelWeightsvariants and clear state uniformly.quantized_gemma3(per-layerOption<(Tensor, Tensor)>), the explicit clear also frees the cached tensors between requests instead of keeping the last conversation's KV alive until the next forward.The implementation copies the exact pattern of
quantized_llama::ModelWeights::clear_kv_cache(iterate layers, clear the K/V pair):kv_cache = Nonefor gemma3,KvCache::reset()for phi3/glm4 (which usecandle_nn::kv_cache::KvCache).No behavior change for existing callers — the method is additive. Same pattern previously merged for the siblings in #3536 (quantized llama/qwen2) and #3189 (quantized qwen3).
Related CPU findings from the same embedding work, filed separately: #3707 (no batch path in the quantized matmul — prefill runs at decode speed) and #3708 (no lazy/mmap loading for GGUF).