Skip to content

Add clear_kv_cache to quantized_llama and quantized_qwen2 - #3536

Merged
ivarflakstad merged 1 commit into
huggingface:mainfrom
Sok205:quantized-clear-kv-cache
May 16, 2026
Merged

ivarflakstad merged 1 commit into
huggingface:mainfrom
Sok205:quantized-clear-kv-cache

Conversation

@Sok205

@Sok205 Sok205 commented May 16, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds clear_kv_cache(&mut self) to ModelWeights in both
candle-transformers/src/models/quantized_llama.rs and
candle-transformers/src/models/quantized_qwen2.rs, mirroring the
existing implementations in qwen3.rs, gemma3.rs, phi.rs,
smollm3.rs, and several others.

pub fn clear_kv_cache(&mut self) {
    for layer in self.layers.iter_mut() {
        layer.kv_cache = None;
    }
}

Motivation

Callers reusing a quantized model across independent conversations
currently have no way to drop cached attention state. Their non-quantized
counterparts already expose this method; this brings the two quantized
variants to parity.

Test plan

  • cargo fmt
  • cargo build -p candle-transformers --lib
  • cargo clippy -p candle-transformers --lib -- -D warnings

Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and
several other transformer models. Lets callers free cached attention
state between independent conversations without recreating the model.

@ivarflakstad ivarflakstad left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks 👌

@ivarflakstad
ivarflakstad merged commit 525b81e into huggingface:main May 16, 2026
11 checks passed
Sok205 added a commit to Sok205/local_vibe that referenced this pull request May 16, 2026
huggingface/candle#3536 (clear_kv_cache on quantized_llama and
quantized_qwen2) was merged upstream. The headless Metal device patch
the fork also carried is already in upstream main. Pin to the merge
commit until a tagged release that includes the change is available.
Sok205 added a commit to Sok205/local_vibe that referenced this pull request May 16, 2026
huggingface/candle#3536 (clear_kv_cache on quantized_llama and
quantized_qwen2) was merged upstream. The headless Metal device patch
the fork also carried is already in upstream main. Pin to the merge
commit until a tagged release that includes the change is available.
jaw-sh added a commit to jaw-sh/candle that referenced this pull request May 19, 2026
The published Qwen3-VL configs set rope_scaling.mrope_interleaved=true
with a [T, H, W] section split summing to head_dim/2. The previous
rotary embedding ignored that block and ran a single 1D counter over
the whole sequence, so vision tokens carried no row/column/temporal
distinction and the model generated text uncorrelated with the image.

Also flips an inverted condition in Qwen3VLModel::forward that built
the causal mask for decode and skipped it for prefill, mirroring
qwen2.rs. With M-RoPE wired up but the mask still inverted, the
model emitted <|im_end|> as its first token after prefill.

RotaryEmbedding now takes optional [T, H, W] band sizes; an empty
section preserves the 1D behaviour byte-for-byte. Qwen3VLModel::forward
takes position_ids of shape (3, batch, seq_len) and a separate
cache_position, since M-RoPE advances positions by max(grid_h, grid_w)
across an image span while the cache grows by the placeholder count.
compute_mrope_position_ids and single_token_position_ids are exported
for callers. The per-layer KV cache switches from KvCache to
ConcatKvCache to match qwen3.rs, and clear_kv_cache is added in the
spirit of huggingface#3536; it takes &self because the existing attention wraps
the cache in Arc<Mutex<>> for the deepstack injection path.
jaw-sh added a commit to jaw-sh/candle that referenced this pull request May 19, 2026
The published Qwen3-VL configs set rope_scaling.mrope_interleaved=true
with a [T, H, W] section split summing to head_dim/2. The previous
rotary embedding ignored that block and ran a single 1D counter over
the whole sequence, so vision tokens carried no row/column/temporal
distinction and the model generated text uncorrelated with the image.

Also flips an inverted condition in Qwen3VLModel::forward that built
the causal mask for decode and skipped it for prefill, mirroring
qwen2.rs. With M-RoPE wired up but the mask still inverted, the
model emitted <|im_end|> as its first token after prefill.

RotaryEmbedding now takes optional [T, H, W] band sizes; an empty
section preserves the 1D behaviour byte-for-byte. Qwen3VLModel::forward
takes position_ids of shape (3, batch, seq_len) and a separate
cache_position, since M-RoPE advances positions by max(grid_h, grid_w)
across an image span while the cache grows by the placeholder count.
compute_mrope_position_ids and single_token_position_ids are exported
for callers. The per-layer KV cache switches from KvCache to
ConcatKvCache to match qwen3.rs, and clear_kv_cache is added in the
spirit of huggingface#3536.
FerrisMind pushed a commit to FerrisMind/candle that referenced this pull request Jun 23, 2026
…e#3536)

Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and
several other transformer models. Lets callers free cached attention
state between independent conversations without recreating the model.
SorenDreano pushed a commit to SorenDreano/candle that referenced this pull request Jul 9, 2026
…e#3536)

Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and
several other transformer models. Lets callers free cached attention
state between independent conversations without recreating the model.
Abhinav5132 pushed a commit to Abhinav5132/candle that referenced this pull request Jul 24, 2026
…e#3536)

Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and
several other transformer models. Lets callers free cached attention
state between independent conversations without recreating the model.
achmadk pushed a commit to achmadk/candle that referenced this pull request Sep 13, 2026
…e#3536)

Mirrors the pattern already used in qwen3, gemma3, phi, smollm3, and
several other transformer models. Lets callers free cached attention
state between independent conversations without recreating the model.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants