Describe the solution you'd like
Seems like there is no support in SFT/LORA/RL of gemma4 model with multi-modal inputs (text+image for example). There is support for gemma3, I would appriciate if you can supply such support and pointers
Checklist
- [ V] I have searched the existing issues for similar feature requests.
- [V ] This is not a support question (please use the "bug template" for that).
Describe the solution you'd like
Seems like there is no support in SFT/LORA/RL of gemma4 model with multi-modal inputs (text+image for example). There is support for gemma3, I would appriciate if you can supply such support and pointers
Checklist