Skip to content

Python: [Feature]: Support prompt cache breakpoints for GPT 5.6 models #7157

Description

Description

The GPT-5.6 models support a new prompt_cache_breakpoint parameter that explicitly sets breakpoints for where prompt prefixes should be saved to the cache. This is important for the newer models, as cache writes now cost.
https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-breakpoints

Code Sample

Perhaps something like this.

Content.from_text(
  "Text input",
  additional_properties={
    "prompt_cache_breakpoint": {
      "mode": "explicit"
    }
  }
)

Language/SDK

Python

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Jul 16, 2026
  2. changed the title [-][Feature]: Support prompt cache breakpoints for GPT 5.6 models[/-] [+]Python: [Feature]: Support prompt cache breakpoints for GPT 5.6 models[/+] on Jul 16, 2026
  3. Mordris commented on Jul 16, 2026

    @Mordris
    Contributor

    I'd like to pick this up if no one's already on it.

    Before opening a PR I looked at how the OpenAI clients assemble requests, and there are a couple of design choices worth settling first.

    Block-level prompt_cache_breakpoint: content parts are built in _prepare_content_for_openai, which doesn't currently carry arbitrary Content metadata onto the outgoing part, so the additional_properties={"prompt_cache_breakpoint": {...}} form from the description would be dropped today. There's already a precedent for the pattern that would make it work: detail, file_id, and filename are read from content.additional_properties and set on the part. The low-footprint option is to read prompt_cache_breakpoint the same way and attach it to the blocks each API accepts it on (text, image, audio and file parts for Chat Completions; input_text/input_image/input_file for Responses). The alternative is a first-class field on Content. Do you have a preference between the additional_properties route and a typed field?

    Request-level prompt_cache_options ({"mode": "implicit" | "explicit", "ttl": "30m"}): this reads as a natural sibling to the existing prompt_cache_key / prompt_cache_retention on OpenAIChatOptions, and _prepare_options already forwards options through, so it should be a small addition.

    SDK floor: the typed params landed in openai==2.45.0, and the package currently pins openai>=2.25.0,<3, so the request-level option would need the floor bumped to >=2.45.0.

    I'd cover both the Responses and Chat Completions clients, add unit tests asserting the params reach the request payload, and validate against the gpt-5.6 models. Happy to build whatever shape you land on.

  4. added
    agentsUsage: [Issues, PRs], Target: Single agent
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Jul 20, 2026
  5. jguizarj commented on Jul 24, 2026

    @jguizarj

    Hi everyone! Are there any plans to add support for prompt cache breakpoints in .NET?

  6. Mordris commented on Jul 25, 2026

    @Mordris
    Contributor

    Hi Jonathan Guizar (@jguizarj) — I dug into this fairly thoroughly (I did the Python side in #7163), so let me share what I found. Two parts: the honest state of play, and a working recipe you can use right now.

    State of play. There's no typed prompt_cache_options / prompt_cache_breakpoint in .NET yet — and, unlike Python, that gap is in the OpenAI .NET SDK, not in Agent Framework. Agent Framework doesn't own OpenAI request-building on .NET; it adapts the OpenAI SDK client into an IChatClient via Microsoft.Extensions.AI.OpenAI. The OpenAI SDK (latest 2.12.0) exposes no prompt-cache-breakpoint API on either Chat Completions or Responses; the closest tracked item, openai/openai-dotnet#853, is cache retention only and is labelled "blocked: upstream." So there's nothing Agent Framework can cleanly surface until the SDK does.

    But you can use GPT‑5.6 explicit caching today through the OpenAI SDK's JsonPatch escape hatch, and it flows correctly through Agent Framework. I verified this end‑to‑end against gpt-5.6-luna — the second turn reports cached input tokens — via both IChatClient and AIAgent, non‑streaming and streaming, and the request also serializes correctly through AzureOpenAIClient (same underlying ChatClient).

    ⚠️ JsonPatch is [Experimental("SCME0001")] ("subject to change") and gives you no compile‑time validation of the option shape. Treat this as a bridge, not a long‑term API.

    Request‑level prompt_cache_options — via ChatOptions.RawRepresentationFactory:

    #pragma warning disable SCME0001 // JsonPatch is experimental
    using Microsoft.Agents.AI;
    using Microsoft.Extensions.AI;
    using OpenAI.Chat;
    
    var options = new ChatOptions
    {
        RawRepresentationFactory = _ =>
        {
            var o = new ChatCompletionOptions();
            o.Patch.Set("$.prompt_cache_options.mode"u8, "explicit");
            return o;
        },
    };

    Per‑part prompt_cache_breakpoint — attach a native content part (carrying its own Patch) via AIContent.RawRepresentation:

    var prefix = ChatMessageContentPart.CreateTextPart(longReusablePrefix);
    prefix.Patch.Set("$.prompt_cache_breakpoint.mode"u8, "explicit");
    
    // Chat Completions gotcha: a lone text part is serialized as a plain string and the
    // breakpoint is silently dropped. Keep the content array-form — e.g. include the
    // question as a second part:
    var nativeMsg = new UserChatMessage(new[]
    {
        prefix,
        ChatMessageContentPart.CreateTextPart(question),
    });
    var messages = new List<ChatMessage>
    {
        new(ChatRole.User, []) { RawRepresentation = nativeMsg },
    };
    
    var agent = openAIClient.GetChatClient("gpt-5.6-luna").AsIChatClient().AsAIAgent();
    var result = await agent.RunAsync(messages, options: new ChatClientAgentRunOptions(options));

    On the Responses API this is cleaner: a single ResponseContentPart with the breakpoint works without the array‑form workaround (Responses always sends content as an array), and the request‑level option goes on CreateResponseOptions.Patch the same way.

    Toward first‑class support. The right lever is a focused feature request on openai-dotnet for prompt_cache_options + prompt_cache_breakpoint (explicit mode) on the request‑options and content‑part types — separate from that retention issue, which only covers retention. Once the SDK exposes typed properties, Microsoft.Extensions.AI.OpenAI (and then Agent Framework) can surface them ergonomically, matching what shipped for Python. I'd be glad to help move that along.

    (Whether Agent Framework adds first‑class .NET support is of course a maintainer call — I'm just sharing what's technically possible today and where the real blocker sits.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

agentsUsage: [Issues, PRs], Target: Single agentpythonUsage: [Issues, PRs], Target: Python

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions