Repository navigation
Conversation
JumpLink
marked this pull request as ready for review
July 28, 2026 17:48
Add a `vision` option to getOpenAIConfig / getOpenAILLMConfig that is
forwarded onto the OpenAI llmConfig (`OAIClientOptions.vision`). When a
caller knows the target model has no vision support, it can set
`vision: false` so the chat client strips image content before sending,
avoiding hard provider errors ("model is not a multimodal model" / "No
endpoints found that support image input").
Defaults to undefined, so existing behavior is unchanged. The image
stripping itself is implemented in @librechat/agents (see
LibreChat-AI/agents#257), which this option drives.
Covered by getOpenAILLMConfig unit tests (vision true/false/absent).
JumpLink
force-pushed
the
feat/vision-capability
branch
from
August 14, 2026 09:38
dc1ca7c to
76c8bb2
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on
LibreChat-AI/agents#257 (the agents-side change that consumes this flag). Draft until that lands - this option is a no-op without it.
Problem
When an agent routes image content to a model with no vision support, OpenAI-compatible providers reject the whole request:
model is not a multimodal modelNo endpoints found that support image inputThis happens in practice when a tool returns image content (e.g. an image-generation tool whose base64 artifact is fed back as context) and the active model is text-only. Today there is no way to tell the OpenAI client "this model can't take images, drop them."
Change
Add an optional
visionflag togetOpenAIConfig/getOpenAILLMConfig, forwarded onto the OpenAIllmConfig(OAIClientOptions.vision). When a caller setsvision: false, the chat client strips image content before sending.OpenAIConfigOptions.vision?: boolean(public input) andOAIClientOptions.vision?: boolean(the client options).getOpenAILLMConfigsetsllmConfig.visionwhen the flag is provided; otherwise it is omitted.Defaults to
undefined, so existing behavior is unchanged unless a caller opts in.Consumption
The image stripping itself lives in
@librechat/agents(LibreChat-AI/agents#257):ChatOpenAI/AzureChatOpenAI/ChatDeepSeek/ChatXAIreadvisionfrom their constructor fields and stripimage_urlparts before streaming. ThellmConfigproduced here becomes those constructor fields, so this flag drives that behavior.The capability source (a model spec / agent
visionboolean →options.vision) is intentionally left to the caller; this PR is the config-layer plumbing.Tests
getOpenAILLMConfigunit tests covervisiontrue / false / absent.Testing status
Built and unit-tested at the
@librechat/apipackage level (typecheck clean, 85/85 inllm.spec.ts). Not yet end-to-end integration tested against agents#257, since that change is unpublished - hence draft.