Repository navigation
Conversation
llama.cpp reports a model's inputs as architecture.input_modalities, and builds before October 2026 only as modalities.vision in /props. Standard discovery read neither, and both parsers rejected lists that include audio or video, so vision models were treated as text-only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e top Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Tau treated llama.cpp vision models as text-only, so the read tool never sent them images.
Standard discovery did not read the field where llama.cpp reports a model's inputs, and older llama.cpp builds report vision only in
/props.This change reads both reports, and it stops one unsupported entry, such as
audio, from hiding image support.Related to #602, which added the llama.cpp backend.
What Changed
llama.cpp reports inputs in two places, depending on its age.
Since llama.cpp #29987 (2026-10-05), every
/v1/modelsentry hasarchitecture.input_modalities.Older builds report vision only as
modalities.visionin/props, which describes the one loaded model of a standard server.reported_input_modalitiesinrouter.pyis now the single parser for a/modelsentry. It readsarchitecture.input_modalities, and falls back to a top-levelinput_modalitiesormodalitieslist. Router mode and standard discovery both use it.audioandvideoentries and drops them, because Tau sends neither. Before, a list such as["text", "image", "audio"]was rejected as a whole. Unknown values still make the report untrusted, as the existing test requires./propswhen the server lists exactly one model and that model reports no inputs. A failing or malformed/propsleaves the model as it was, so discovery never fails because of it.catalog.toml.Testing
I added unit tests for both parsers and discovery tests for the new paths, and I ran the change against two real llama.cpp servers.
uv run pytest tests/test_llama_cpp_extension.py: 66 passed. New cases: architecture modalities with audio, the/propsfallback for vision and text-only models, a failing/props,/modelstaking precedence over/props, and no fallback for several models.uv run pytest: 2069 passed and 1 failed. The failure istest_tui_tree_labels_filter_timestamps_and_clear, which expects UTC. It fails on my UTC+8 machine and passes withTZ=UTC, so it is unrelated to this change.uv run ruff check .,uv run ruff format --check .anduv run mypypass.--mmproj: before, a fresh Tau treated the model as text-only. With this branch,tau -p --provider llama.cppread an invoice image with the read tool and answered with the correct total and shop name, with nocatalog.tomlentry. This build has noarchitecturefield, so the/propsfallback did the work.architecture.input_modalities(only covered by unit tests), router mode against a live server, and the Hugo docs build (Hugo is not installed here).Risks
The risk is low, because the change only adds information that the server reports directly.
Tau still never guesses capabilities from a model name.
vision: truein/propsfor a model that cannot see would now receive images. That would be a server bug, and the same is true of the/modelsreport./propsfallback adds one request to standard discovery, only when a single model reports no inputs.