Skip to content

feat(models): port Llama 3.2 Vision (cross-attention vision + Llama-3 text) #527

Description

@inureyes

Summary

Add Llama 3.2 Vision support to mlxcel. Its upstream model_type is mllama. mlxcel currently has no arm for it in src/models/detection.rs, so a mllama checkpoint errors "Unsupported model type".

What it is

Meta Llama-3.2-Vision; a Llama-3 text backbone with gated cross-attention layers attending to a ViT image encoder's features.

In-tree reuse

Llama text backbone src/models/llama3.rs.

Scope

  • Add the ModelType variant + detection.rs arm for mllama.
  • ViT vision encoder + cross-attention adapter layers + image processor.
  • Register across the arch/metadata tables; update docs/supported-models.md; add tests + validate on a real checkpoint.

Effort: medium-high

Metadata

Metadata

Assignees

Labels

area:modelsModel architectures, weights, loading, metadatapriority:mediumMedium prioritystatus:doneCompletedtype:enhancementNew features, capabilities, or significant additions

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions