Skip to content

binding-llm: canonical section stream, with OpenAI as the reference dialect #2634

Description

@jfallows

Depends on: #2633

Summary

llm(server) decodes the inbound dialect into the canonical block sequence; llm(client) encodes it into the backend dialect. OpenAI Chat Completions is the first dialect, so that the framework is shaped by one concrete mapping before the others are ported.

Design

Section types (the content-block sections are rows 1-14; request-level and block-level controls that are not content blocks are carried as dialect extension fields, see #2638):

# Section openai anthropic ollama
1 system-instruction system/developer msg top-level system system msg
2 user-text yes yes yes
3 user-image image_url image images[]
4 user-document file document no
5 user-audio input_audio no no
6 user-video no no no
7 assistant-text yes yes yes
8 assistant-reasoning no (Chat Completions) thinking, redacted_thinking think param, no replay block
9 assistant-refusal refusal no no
10 tool-definition tools[] tools[] tools[]
11 tool-call tool_calls tool_use tool_calls
12 tool-result role: tool tool_result role: tool
13 server-tool-result no server_tool_use, web_search_tool_result, web_fetch, code_execution_result no
14 citations no on text blocks no

A section outside the taxonomy is unknown, and that is all unknown means.

  • The block, not the dialect message, is the unit of the canonical stream: an anthropic user message holding [text, image, tool_result] becomes three blocks sharing one message index.
  • llm(client) reassembles messages from message and derives role from section type and target dialect.
  • The section type is carried on every block, so llm(client) can derive the target dialect's role and message structure from it regardless of the source dialect's wire shape.
  • Replaces LlmCanonicalEvent / LlmCanonicalEmitter / LlmCanonicalEncodeSink for the request side, and the response side uses the same block shape.

Needs a decision before implementation (spike first)

Decode and encode must not buffer the body, but some dialect reassembly needs reordering: an anthropic top-level system that arrives after messages must precede them in an openai messages[], and tools[] can be the last JSON key. Either the encode constrains what it will accept across dialects, or it buffers a bounded amount. The design says nothing buffers, so this may contradict it. Report what the spike finds and adjust the design rather than working around it.

Acceptance criteria (test-first)

  • .rpt scripts and unit tests for openai: each section type above, mixed-content messages, fragmented large blocks (re-use the existing 10k and 100k scenarios).
  • Round trip openai to openai reproduces an equivalent request.
  • Spike outcome on ordering/buffering recorded on the issue.
  • Existing OpenAI ITs still pass or are migrated with the reason noted.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions