Repository navigation
Add clear_kv_cache to quantized_llama and quantized_qwen2 - #3536
Merged
Merged
Conversation
Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and several other transformer models. Lets callers free cached attention state between independent conversations without recreating the model.
Sok205
added a commit
to Sok205/local_vibe
that referenced
this pull request
May 16, 2026
huggingface/candle#3536 (clear_kv_cache on quantized_llama and quantized_qwen2) was merged upstream. The headless Metal device patch the fork also carried is already in upstream main. Pin to the merge commit until a tagged release that includes the change is available.
Sok205
added a commit
to Sok205/local_vibe
that referenced
this pull request
May 16, 2026
huggingface/candle#3536 (clear_kv_cache on quantized_llama and quantized_qwen2) was merged upstream. The headless Metal device patch the fork also carried is already in upstream main. Pin to the merge commit until a tagged release that includes the change is available.
jaw-sh
added a commit
to jaw-sh/candle
that referenced
this pull request
May 19, 2026
The published Qwen3-VL configs set rope_scaling.mrope_interleaved=true with a [T, H, W] section split summing to head_dim/2. The previous rotary embedding ignored that block and ran a single 1D counter over the whole sequence, so vision tokens carried no row/column/temporal distinction and the model generated text uncorrelated with the image. Also flips an inverted condition in Qwen3VLModel::forward that built the causal mask for decode and skipped it for prefill, mirroring qwen2.rs. With M-RoPE wired up but the mask still inverted, the model emitted <|im_end|> as its first token after prefill. RotaryEmbedding now takes optional [T, H, W] band sizes; an empty section preserves the 1D behaviour byte-for-byte. Qwen3VLModel::forward takes position_ids of shape (3, batch, seq_len) and a separate cache_position, since M-RoPE advances positions by max(grid_h, grid_w) across an image span while the cache grows by the placeholder count. compute_mrope_position_ids and single_token_position_ids are exported for callers. The per-layer KV cache switches from KvCache to ConcatKvCache to match qwen3.rs, and clear_kv_cache is added in the spirit of huggingface#3536; it takes &self because the existing attention wraps the cache in Arc<Mutex<>> for the deepstack injection path.
jaw-sh
added a commit
to jaw-sh/candle
that referenced
this pull request
May 19, 2026
The published Qwen3-VL configs set rope_scaling.mrope_interleaved=true with a [T, H, W] section split summing to head_dim/2. The previous rotary embedding ignored that block and ran a single 1D counter over the whole sequence, so vision tokens carried no row/column/temporal distinction and the model generated text uncorrelated with the image. Also flips an inverted condition in Qwen3VLModel::forward that built the causal mask for decode and skipped it for prefill, mirroring qwen2.rs. With M-RoPE wired up but the mask still inverted, the model emitted <|im_end|> as its first token after prefill. RotaryEmbedding now takes optional [T, H, W] band sizes; an empty section preserves the 1D behaviour byte-for-byte. Qwen3VLModel::forward takes position_ids of shape (3, batch, seq_len) and a separate cache_position, since M-RoPE advances positions by max(grid_h, grid_w) across an image span while the cache grows by the placeholder count. compute_mrope_position_ids and single_token_position_ids are exported for callers. The per-layer KV cache switches from KvCache to ConcatKvCache to match qwen3.rs, and clear_kv_cache is added in the spirit of huggingface#3536.
4 tasks done
FerrisMind
pushed a commit
to FerrisMind/candle
that referenced
this pull request
Jun 23, 2026
…e#3536) Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and several other transformer models. Lets callers free cached attention state between independent conversations without recreating the model.
SorenDreano
pushed a commit
to SorenDreano/candle
that referenced
this pull request
Jul 9, 2026
…e#3536) Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and several other transformer models. Lets callers free cached attention state between independent conversations without recreating the model.
Abhinav5132
pushed a commit
to Abhinav5132/candle
that referenced
this pull request
Jul 24, 2026
…e#3536) Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and several other transformer models. Lets callers free cached attention state between independent conversations without recreating the model.
achmadk
pushed a commit
to achmadk/candle
that referenced
this pull request
Sep 13, 2026
…e#3536) Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and several other transformer models. Lets callers free cached attention state between independent conversations without recreating the model.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
clear_kv_cache(&mut self)toModelWeightsin bothcandle-transformers/src/models/quantized_llama.rsandcandle-transformers/src/models/quantized_qwen2.rs, mirroring theexisting implementations in
qwen3.rs,gemma3.rs,phi.rs,smollm3.rs, and several others.Motivation
Callers reusing a quantized model across independent conversations
currently have no way to drop cached attention state. Their non-quantized
counterparts already expose this method; this brings the two quantized
variants to parity.
Test plan
cargo fmtcargo build -p candle-transformers --libcargo clippy -p candle-transformers --lib -- -D warnings