Skip to content

feat: optional vision flag in OpenAI LLM config - #13860

Open
JumpLink wants to merge 1 commit into
LibreChat-AI:mainfrom
faktenforum:feat/vision-capability
Open

JumpLink wants to merge 1 commit into
LibreChat-AI:mainfrom
faktenforum:feat/vision-capability

Conversation

@JumpLink

Copy link
Copy Markdown
Contributor

Depends on

LibreChat-AI/agents#257 (the agents-side change that consumes this flag). Draft until that lands - this option is a no-op without it.

Problem

When an agent routes image content to a model with no vision support, OpenAI-compatible providers reject the whole request:

  • model is not a multimodal model
  • No endpoints found that support image input

This happens in practice when a tool returns image content (e.g. an image-generation tool whose base64 artifact is fed back as context) and the active model is text-only. Today there is no way to tell the OpenAI client "this model can't take images, drop them."

Change

Add an optional vision flag to getOpenAIConfig / getOpenAILLMConfig, forwarded onto the OpenAI llmConfig (OAIClientOptions.vision). When a caller sets vision: false, the chat client strips image content before sending.

  • OpenAIConfigOptions.vision?: boolean (public input) and OAIClientOptions.vision?: boolean (the client options).
  • getOpenAILLMConfig sets llmConfig.vision when the flag is provided; otherwise it is omitted.
  • Scoped to the OpenAI branch only (the ChatOpenAI / Azure / DeepSeek / xAI family that agents#257 covers). Anthropic/Google are untouched.

Defaults to undefined, so existing behavior is unchanged unless a caller opts in.

Consumption

The image stripping itself lives in @librechat/agents (LibreChat-AI/agents#257): ChatOpenAI/AzureChatOpenAI/ChatDeepSeek/ChatXAI read vision from their constructor fields and strip image_url parts before streaming. The llmConfig produced here becomes those constructor fields, so this flag drives that behavior.

The capability source (a model spec / agent vision boolean → options.vision) is intentionally left to the caller; this PR is the config-layer plumbing.

Tests

getOpenAILLMConfig unit tests cover vision true / false / absent.

Testing status

Built and unit-tested at the @librechat/api package level (typecheck clean, 85/85 in llm.spec.ts). Not yet end-to-end integration tested against agents#257, since that change is unpublished - hence draft.

@JumpLink
JumpLink marked this pull request as ready for review July 28, 2026 17:48
Add a `vision` option to getOpenAIConfig / getOpenAILLMConfig that is
forwarded onto the OpenAI llmConfig (`OAIClientOptions.vision`). When a
caller knows the target model has no vision support, it can set
`vision: false` so the chat client strips image content before sending,
avoiding hard provider errors ("model is not a multimodal model" / "No
endpoints found that support image input").

Defaults to undefined, so existing behavior is unchanged. The image
stripping itself is implemented in @librechat/agents (see
LibreChat-AI/agents#257), which this option drives.

Covered by getOpenAILLMConfig unit tests (vision true/false/absent).
@JumpLink
JumpLink force-pushed the feat/vision-capability branch from dc1ca7c to 76c8bb2 Compare August 14, 2026 09:38
@danny-avila danny-avila added area/LLM provider integration and observability 🗺️ LLM Provider Config codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) and removed area/LLM provider integration and observability labels Sep 24, 2026
@codegraph-librechat codegraph-librechat Bot added the 🗺️ Backend Infra codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) label Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🗺️ Backend Infra codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) 🗺️ LLM Provider Config codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants