Skip to content

fix(llama-cpp): detect image input from /models architecture and /props - #757

Open
osolmaz wants to merge 2 commits into
huggingface:mainfrom
osolmaz:fix/llama-cpp-architecture-modalities
Open

osolmaz wants to merge 2 commits into
huggingface:mainfrom
osolmaz:fix/llama-cpp-architecture-modalities

Conversation

@osolmaz

@osolmaz osolmaz commented Oct 6, 2026

Copy link
Copy Markdown
Member

Summary

Tau treated llama.cpp vision models as text-only, so the read tool never sent them images.
Standard discovery did not read the field where llama.cpp reports a model's inputs, and older llama.cpp builds report vision only in /props.
This change reads both reports, and it stops one unsupported entry, such as audio, from hiding image support.
Related to #602, which added the llama.cpp backend.

What Changed

llama.cpp reports inputs in two places, depending on its age.
Since llama.cpp #29987 (2026-10-05), every /v1/models entry has architecture.input_modalities.
Older builds report vision only as modalities.vision in /props, which describes the one loaded model of a standard server.

  • reported_input_modalities in router.py is now the single parser for a /models entry. It reads architecture.input_modalities, and falls back to a top-level input_modalities or modalities list. Router mode and standard discovery both use it.
  • The parser accepts llama.cpp's audio and video entries and drops them, because Tau sends neither. Before, a list such as ["text", "image", "audio"] was rejected as a whole. Unknown values still make the report untrusted, as the existing test requires.
  • Standard discovery now falls back to /props when the server lists exactly one model and that model reports no inputs. A failing or malformed /props leaves the model as it was, so discovery never fails because of it.
  • The local inference guide explains how image input is detected, and that a server reporting nothing should be declared in catalog.toml.

Testing

I added unit tests for both parsers and discovery tests for the new paths, and I ran the change against two real llama.cpp servers.

  • uv run pytest tests/test_llama_cpp_extension.py: 66 passed. New cases: architecture modalities with audio, the /props fallback for vision and text-only models, a failing /props, /models taking precedence over /props, and no fallback for several models.
  • uv run pytest: 2069 passed and 1 failed. The failure is test_tui_tree_labels_filter_timestamps_and_clear, which expects UTC. It fails on my UTC+8 machine and passes with TZ=UTC, so it is unrelated to this change.
  • uv run ruff check ., uv run ruff format --check . and uv run mypy pass.
  • Live, official llama.cpp b10711 with Qwen3-VL-8B and its --mmproj: before, a fresh Tau treated the model as text-only. With this branch, tau -p --provider llama.cpp read an invoice image with the read tool and answered with the correct total and shop name, with no catalog.toml entry. This build has no architecture field, so the /props fallback did the work.
  • Not tested: a llama.cpp build that already reports architecture.input_modalities (only covered by unit tests), router mode against a live server, and the Hugo docs build (Hugo is not installed here).

Risks

The risk is low, because the change only adds information that the server reports directly.
Tau still never guesses capabilities from a model name.

  • A server that reports vision: true in /props for a model that cannot see would now receive images. That would be a server bug, and the same is true of the /models report.
  • The /props fallback adds one request to standard discovery, only when a single model reports no inputs.

llama.cpp reports a model's inputs as architecture.input_modalities, and
builds before October 2026 only as modalities.vision in /props. Standard
discovery read neither, and both parsers rejected lists that include audio
or video, so vision models were treated as text-only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@osolmaz
osolmaz requested a review from alejandro-ao as a code owner October 6, 2026 10:18
…e top

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant