Problem
In the legacy OpenAI-compatible endpoint /v1/chat/completions, image content works correctly when it is sent in a normal user message.
However, if a tool message returns image content (for example structured array content containing image_url or file-based multimodal parts), the Codex path fails and the client receives:
empty_stream: upstream stream closed before first payload
Current behavior
user message with image input: works
tool message with image content: fails
This suggests the issue is in the translator path for tool-role content, not in general multimodal handling.
Expected behavior
Tool-role content should preserve multimodal payloads in the same way user-role content does, as long as the downstream Responses payload remains valid.
Suspected area
internal/translator/codex/openai/chat-completions/codex_openai_request.go
The translator likely needs to support structured tool content arrays and correctly map multimodal parts for function_call_output.output.
Required behavior
Please update the translator so that:
- structured tool content arrays can preserve supported multimodal parts such as
text, image_url, and valid file parts
input_image is emitted only when a real image source exists (url or file_id)
input_file is emitted only when a real file source exists (file_id, file_data, or file_url)
- unsupported or incomplete parts are not silently dropped in a lossy way; they should fall back to raw text or serialized JSON
function_call_output.output should remain present even when the original tool content is not a string/array (for example null or other JSON values)
Reproduction hints
Observed symptom:
empty_stream: upstream stream closed before first payload
Repro pattern:
- send request to
/v1/chat/completions
- let the model call a tool
- return tool content containing image data or multimodal structured content
- observe upstream stream closes before first payload
Regression coverage
Please add tests covering:
- tool content with
image_url
- tool content with valid file parts
- filename-only file parts should not emit invalid
input_file
- image parts without source should not emit invalid
input_image
content: null
- unknown tool content part types should not be silently dropped
Related context
- Related PR attempt:
#2349
Problem
In the legacy OpenAI-compatible endpoint
/v1/chat/completions, image content works correctly when it is sent in a normalusermessage.However, if a tool message returns image content (for example structured array content containing
image_urlor file-based multimodal parts), the Codex path fails and the client receives:Current behavior
usermessage with image input: workstoolmessage with image content: failsThis suggests the issue is in the translator path for tool-role content, not in general multimodal handling.
Expected behavior
Tool-role content should preserve multimodal payloads in the same way user-role content does, as long as the downstream Responses payload remains valid.
Suspected area
internal/translator/codex/openai/chat-completions/codex_openai_request.goThe translator likely needs to support structured tool content arrays and correctly map multimodal parts for
function_call_output.output.Required behavior
Please update the translator so that:
text,image_url, and validfilepartsinput_imageis emitted only when a real image source exists (urlorfile_id)input_fileis emitted only when a real file source exists (file_id,file_data, orfile_url)function_call_output.outputshould remain present even when the original toolcontentis not a string/array (for examplenullor other JSON values)Reproduction hints
Observed symptom:
Repro pattern:
/v1/chat/completionsRegression coverage
Please add tests covering:
image_urlinput_fileinput_imagecontent: nullRelated context
#2349