Summary
Add Llama 3.2 Vision support to mlxcel. Its upstream model_type is mllama. mlxcel currently has no arm for it in src/models/detection.rs, so a mllama checkpoint errors "Unsupported model type".
What it is
Meta Llama-3.2-Vision; a Llama-3 text backbone with gated cross-attention layers attending to a ViT image encoder's features.
In-tree reuse
Llama text backbone src/models/llama3.rs.
Scope
- Add the
ModelType variant + detection.rs arm for mllama.
- ViT vision encoder + cross-attention adapter layers + image processor.
- Register across the arch/metadata tables; update
docs/supported-models.md; add tests + validate on a real checkpoint.
Effort: medium-high
Summary
Add Llama 3.2 Vision support to mlxcel. Its upstream model_type is
mllama. mlxcel currently has no arm for it insrc/models/detection.rs, so amllamacheckpoint errors "Unsupported model type".What it is
Meta Llama-3.2-Vision; a Llama-3 text backbone with gated cross-attention layers attending to a ViT image encoder's features.
In-tree reuse
Llama text backbone
src/models/llama3.rs.Scope
ModelTypevariant +detection.rsarm formllama.docs/supported-models.md; add tests + validate on a real checkpoint.Effort: medium-high