Repository navigation
Conversation
Multimodal callers need to build a mixed embedding sequence - text embeddings with image embeddings spliced in - which cannot be expressed as token ids. gemma already exposes both entry points and paligemma is built on them; gemma3 exposed neither, so Gemma 3's vision variants cannot be wired up outside the crate. forward now delegates to forward_embeds, so there is no behaviour change: the two produce bit-identical logits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
gemma3::Modelonly accepts token ids, and keepsembed_tokensprivate. A multimodal caller needs to build a mixed sequence — text embeddings with image embeddings spliced in — which has no representation as token ids, so Gemma 3's vision variants can't be wired up from outside the crate.gemma.rsalready exposes both entry points (embed_tokens()at line 371,forward_embeds()at line 412) andpaligemma.rsis built on them. This adds the same pair togemma3.forwardnow delegates toforward_embeds, so there's no behaviour change. Unlikegemma,forward_embedsbuilds the attention masks itself rather than taking them as an argument, because gemma3 uses two (full and sliding) and picks per layer — having the caller supply them would leak that detail.Verified the two paths agree exactly (tiny randomly-initialised model, CPU, no downloads):
Note that running that check at all requires #3902 —
gemma3::Model::forwardcurrently fails for every input withslice-set only supports contiguous tensors. This PR compiles and is independent of it, but can't be exercised at runtime until that one lands.