Skip to content

Fix quantized Gemma 3 MLP activation (gelu_pytorch_tanh, not SiLU) - #3747

Open
Sammy-Dabbas wants to merge 1 commit into
huggingface:mainfrom
Sammy-Dabbas:fix/quantized-gemma3-gelu
Open

Sammy-Dabbas wants to merge 1 commit into
huggingface:mainfrom
Sammy-Dabbas:fix/quantized-gemma3-gelu

Conversation

@Sammy-Dabbas

Copy link
Copy Markdown
Contributor

Fixes #3745

The quantized Gemma 3 MLP hardcoded SiLU, but Gemma 3 gates its feed-forward network with gelu_pytorch_tanh. The non-quantized gemma3.rs applies cfg.hidden_activation (gelu_pytorch_tanh for Gemma 3 checkpoints). The quantized model loads from GGUF metadata and has no config field, so it applies the activation directly, the same way quantized_glm4.rs hardcodes its Activation. This routes through the same code path as the non-quantized model (Activation::GeluPytorchTanh, the gelu tanh approximation).

SiLU and gelu_pytorch_tanh behave similarly, so the wrong one does not crash or produce obvious garbage; it just makes the output worse, which is why it went unnoticed. #3326 attempted a sliding-window plus GELU rework but was self-closed by its author, so this is the focused activation-only fix.

@jkantsios

Copy link
Copy Markdown

I was curious about the accuracy/performance impact this would have and produced the following perplexity score changes comparing SiLU and gelu_pytorch.

image image

This definitely seems to confirm the efficacy of this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Gemma models use the wrong activation function

2 participants