Skip to content

Bundled CLI never reads prompt cache in multi-turn headless runs (direct anthropic SDK over the same endpoint caches correctly) #974

Description

@raspin-home

Summary

In a headless claude-agent-sdk run (Python, custom system prompt, MCP
servers configured via .mcp.json), the bundled Claude Code CLI writes
prompt cache on every turn but never reads it — cache_read_input_tokens=0
across multi-turn investigations. On the same endpoint, same model, same
workspace credentials, a 30-line script using the plain anthropic Python
SDK with cache_control: ephemeral on the system block gets a healthy
cache read on the second call.

This makes the regression CLI-side rather than gateway- or API-side, and
the cache_create-on-every-turn behaviour quickly bursts past per-minute
input-token rate limits on multi-turn runs.

Environment

  • claude-agent-sdk versions tested: 0.1.81 (bundled CLI 2.1.139) and
    0.2.82 (bundled CLI 2.1.142) — both broken.
  • Python 3.12 inside a Debian-slim Docker image (Node.js 20 installed for
    the bundled CLI shebang).
  • Model: claude-sonnet-4-6.
  • Endpoint: the Claude on AWS first-party endpoint
    (https://aws-external-anthropic.<region>.api.aws) with an
    anthropic-workspace-id custom header. Same endpoint used by the healthy
    direct-SDK probe below.
  • MCP servers configured via .mcp.json (two HTTP-transport servers).
  • permission_mode="bypassPermissions", strict_mcp_config=True, custom
    system_prompt string (not the claude_code preset), cwd scoped to a
    read-only mount, allowed_tools / disallowed_tools enforced natively.

Symptom — bundled CLI (broken)

Representative log line from a 3-turn run (annotated):

ResultMessage received turns=3 cost_usd=0.187 stop_reason=stop_sequence
usage input=3 output=240 cache_read=0 cache_create=45781
agent completed with error(s)=[] api_error_status=429 subtype=success
  • cache_create_input_tokens ≈ 14–15K per turn, accumulating to
    43–46K across 3 turns.
  • cache_read_input_tokens = 0 across every turn.
  • The third turn trips the workspace's per-minute input-token cap and the
    CLI exits with is_error=true, subtype=success, api_error_status=429.

The "ResultMessage…subtype=success api_error_status=429" combo is
reported via PR #923 — that part is working correctly. The cache miss
is the underlying issue.

Symptom — direct anthropic SDK over the same endpoint (healthy)

[call 1] cache_creation_input_tokens=2722, cache_read_input_tokens=0
[call 2] cache_creation_input_tokens=4,    cache_read_input_tokens=2722

Same ANTHROPIC_BASE_URL, same anthropic-workspace-id header, same
claude-sonnet-4-6 model, same ~11K-char system block carrying
cache_control: {"type": "ephemeral"}.

Conclusion: the gateway / API / workspace / model are all fine. The
bundled CLI is shipping a payload that defeats the cache lookup despite
the prefix appearing identical across turns.

Minimal repro

Two pieces:

1. Healthy reference — direct SDK call, ~50 LOC

import os, time
from anthropic import Anthropic

SYSTEM = ("This is a static filler line used to grow the system prompt past "
          "the cache minimum block size. ") * 80

client = Anthropic(
    api_key=os.environ["ANTHROPIC_API_KEY"],
    base_url=os.environ.get("ANTHROPIC_BASE_URL") or None,
    default_headers={"anthropic-workspace-id": os.environ["ANTHROPIC_WORKSPACE_ID"]}
        if os.environ.get("ANTHROPIC_WORKSPACE_ID") else None,
)

def call(label):
    r = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=64,
        system=[{"type": "text", "text": SYSTEM,
                 "cache_control": {"type": "ephemeral"}}],
        messages=[{"role": "user", "content": "ping"}],
    )
    print(label, r.usage.model_dump())

call("call 1"); time.sleep(1); call("call 2")
# call 2 shows cache_read_input_tokens > 0

2. Broken reference — same workspace, same endpoint, via claude-agent-sdk

from claude_agent_sdk import ClaudeAgentOptions, query
from claude_agent_sdk.types import ResultMessage
import asyncio, os

options = ClaudeAgentOptions(
    system_prompt="<a static ~1.5K-token system prompt>",
    model="claude-sonnet-4-6",
    mcp_servers="/tmp/.mcp.json",       # any two HTTP MCP servers
    strict_mcp_config=True,
    allowed_tools=[...],                # ~16K tokens of tool defs total
    permission_mode="bypassPermissions",
    cwd="/some/read-only/path",
    env={
        "ANTHROPIC_API_KEY": os.environ["ANTHROPIC_API_KEY"],
        "ANTHROPIC_BASE_URL": os.environ["ANTHROPIC_BASE_URL"],
        "ANTHROPIC_CUSTOM_HEADERS":
            f"anthropic-workspace-id: {os.environ['ANTHROPIC_WORKSPACE_ID']}",
    },
    max_turns=50,
)

async def main():
    async for msg in query(prompt="<a small ticket-shaped user message>",
                            options=options):
        if isinstance(msg, ResultMessage):
            print(msg.usage)  # observe cache_read=0 across turns

asyncio.run(main())

The model takes a couple of cheap tool calls and the second/third turn's
usage.cache_read_input_tokens is still 0.

What we ruled out

  • SDK version regression in the v0.2.x line — pinned back to v0.1.81
    (bundled CLI 2.1.139); same broken behaviour.
  • Async MCP loading (v0.2.82's breaking change) — not the cause;
    v0.1.81 exhibits the same symptom.
  • Gateway / workspace / model — direct anthropic SDK over the same
    endpoint shows healthy cache reads.
  • MCP_CONNECTION_NONBLOCKING=0 — env var was set; no observable
    effect (we may have had the env var name wrong, but it didn't move the
    needle either way).

Hypotheses

Things the CLI may be doing differently from the direct SDK call, any of
which would defeat the prefix match:

  1. Per-turn dynamic injection into the system block — date, session
    token, tool registration metadata, or similar non-stable text appearing
    inside the cache_control-marked block.
  2. Tool-array re-ordering between turns once MCP server connections
    settle.
  3. cache_control breakpoint placement drifting between turn 1 and
    turn 2 (e.g., placed on different block indices once more conversation
    accumulates), so it doesn't match what was originally written.
  4. Related to subprocess CLI rejects list-form system_prompt (Anthropic API supports it) #899 (subprocess CLI rejects list-form system_prompt) —
    if the CLI is forced to flatten system to a string, the
    cache_control marker placement may not survive intact.

We can't observe which is true from outside the bundled binary. If anyone
on the SDK side can confirm what the CLI sends as system and tools
between turns when given a custom system_prompt plus MCP servers, that
would settle it quickly.

Impact

  • Multi-turn agentic runs against any workspace with a strict per-minute
    input-token cap will trip 429 within 2–3 turns even on small (<20K-token)
    prefixes.
  • Cost: every turn writes the full prefix to cache fresh, ~3× the expected
    amortised cost.
  • Severity is amplified for the Claude on AWS endpoint where ITPM
    defaults are lower than direct Anthropic accounts.

Asks

  1. Confirm whether dynamic per-call content is being injected into the
    system block in headless mode, and if so, document the env var or
    option to suppress it (analogous to exclude_dynamic_sections on the
    preset).
  2. If not, a way to log the bundled CLI's outbound API request body would
    be sufficient to diagnose from outside.

Happy to share the worker-side and probe scripts in their entirety if it
helps shorten triage.

Activity

  1. alexisoller-cpu commented on May 20, 2026

    @alexisoller-cpu

    Confirmando este bug en producción independientemente.

    Observamos el mismo patrón en headless multi-turn con claude-agent-sdk==0.2.82 (Sonnet 4.6, CLI 2.1.142) lanzados como Server tmux con --dangerously-skip-permissions --print:

    • Cada turno escribe ~14-15K tokens cache_creation_input_tokens y lee 0 tokens cache_read_input_tokens.
    • Aprox. a los 2-3 turnos disparamos 429 por token-rate.
    • Misma key + mismo modelo + misma session usados con anthropic SDK directo (messages.create con cache_control: {type: "ephemeral"} explícito) sí registra cache_read_input_tokens > 0 en los turnos 2+.

    Esto coincide con tu repro al 100%: el caching layer funciona en el endpoint, lo que no lo activa correctamente es el bundled CLI cuando lo invoca el SDK.

    Aprovecho para preguntar: ¿hay un ETA público para fix, o un workaround oficial recomendado mientras tanto (e.g. flag, env var, o downgrade a una versión SDK/CLI donde sí cacheaba)? Para nosotros decide entre seguir migrando a claude-agent-sdk o quedarnos en anthropic SDK directo con tool-loop manual.

    Gracias por el repro tan claro — facilita mucho debug del lado del consumidor.

  2. raspin-home commented on May 21, 2026

    @raspin-home
    Author

    Thanks for the independent confirmation, @alexisoller-cpu - useful to have this off the Claude-on-AWS endpoint, since it rules out gateway-specific behaviour as the cause. The numerical fingerprint matches exactly on our side too.

    It might be worth mentioning that pinning back to claude-agent-sdk==0.1.81 (bundled CLI 2.1.139) had the same cache_read=0. So a simple SDK/CLI downgrade is not the workaround.

    Our current mitigations are:

    • Gating heavy MCP servers per-task so most runs only load the essential tools (smaller tools array → smaller per-turn write).
    • Trimming the tool allowlist down to only what we actually call, in case the SDK filters at the API-request level (still verifying).
    • Requesting an ITPM cap raise on our workspace so multi-turn runs complete despite the cache miss.
  3. raspin-home commented on May 22, 2026

    @raspin-home
    Author

    Quick update with new evidence.

    After @alexisoller-cpu's independent confirmation (same symptom on direct
    Anthropic, not just on our Claude-on-AWS endpoint), I re-ran the probe with
    the Cache Diagnostics beta
    (cache-diagnosis-2026-04-07) attached and threaded
    diagnostics.previous_message_id through the two calls.

    Result on the direct-SDK probe (same endpoint, same workspace, same model
    as the broken worker run):

    [call 1] usage       = {... cache_creation_input_tokens=2722, cache_read_input_tokens=0 ...}
    [call 1] diagnostics = None
    [call 2] usage       = {... cache_creation_input_tokens=4, cache_read_input_tokens=2722 ...}
    [call 2] diagnostics = None
    

    Per the Cache Diagnostics docs, diagnostics: null on call 2 (with a real
    previous_message_id) means "a comparison ran and found no divergence."
    So we now have two independent positive signals on the direct-SDK side:

    1. The cache layer actually delivered cached tokens (cache_read=2722).
    2. Cache Diagnostics formally confirms the request fingerprint is byte-stable
      across turns (diagnostics=null, no cache_miss_reason).

    Together those eliminate "gateway / workspace / model / API caching / our
    own request stability" as possible explanations. The only thing left that
    differs between the healthy direct-SDK case and the broken
    claude-agent-sdk → bundled CLI case is whatever the CLI is doing to its
    outbound request payload between turns.

    (Side note for anyone else who lands here: Cache Diagnostics works over the
    https://aws-external-anthropic.<region>.api.aws endpoint — the doc only
    excludes Bedrock and Vertex, and we confirmed it's honoured on Claude-on-AWS.)

    Concrete ask for SDK-side triage: the official next step the Anthropic
    docs themselves point at for this class of bug is to enable
    cache-diagnosis-2026-04-07 and pass diagnostics.previous_message_id
    between turns. If the bundled CLI's request loop did that on multi-turn
    runs (or if there's an env var / option to surface it), the response's
    cache_miss_reason.type field would directly name the divergent component —
    model_changed / system_changed / tools_changed / messages_changed —
    and the fix follows from there. The hypotheses in the original issue
    (per-turn injection into the system block, tool re-ordering after async MCP
    load, breakpoint placement drift) all map cleanly onto one of those four
    types, so the diagnostic immediately distinguishes them.

    Happy to share the probe script or any worker-side logs.

  4. qing-ant commented on Jun 16, 2026

    @qing-ant
    Contributor

    Thanks for the detailed report. You should be able to unblock yourself immediately by adding this to your options.env:

    env={ ..., "CLAUDE_CODE_ATTRIBUTION_HEADER": "0" }
    

    This would drop a block which invalidates the prompt cache block. We are preparing the fix-forward for an upcoming SDK release so this will no longer be necessary on future releases.

  5. Magic-Man-us commented on Jun 27, 2026

    @Magic-Man-us

    Verified fixed as of bundled CLI 2.1.181 (re-tested on claude-agent-sdk 0.2.110 / CLI 2.1.195): across a multi-turn run cache_read_input_tokens is now non-zero and grows as the prefix accumulates while cache_creation shrinks — prefix caching is working.

    Root cause matches the CLI 2.1.181 CHANGELOG entry — "Fixed prompt caching not reading on custom ANTHROPIC_BASE_URL and on Foundry due to a per-request attestation token changing every turn" — which is exactly the per-turn dynamic prefix change hypothesized in this issue:
    https://github.com/anthropics/claude-code/blob/01f1617f14452ac78bf319cef2236d87c0fe05cb/CHANGELOG.md#L159

    Upgrade claude-agent-sdk to a release bundling CLI ≥ 2.1.181 to pick up the fix.

  6. anirudherabelly commented on Jul 1, 2026

    @anirudherabelly

    @qing-ant, looks like the fix for this issue has been verified by @Magic-Man-us. We can close this bug.

  7. raspin-home commented on Jul 7, 2026

    @raspin-home
    Author

    Confirming the fix on our side — many thanks.

    TL;DR: prefix caching is working again over the Claude-on-AWS first-party endpoint (aws-external-anthropic.<region>.api.aws + anthropic-workspace-id header), and — notably — it's healthy even on the same bundled CLI that originally reproduced the bug, without upgrading. That suggests the fix propagated server/gateway-side for this endpoint, not only via the CLI ≥ 2.1.181 bump.

    To verify I wrote a small CLI-path two-turn probe — driving claude_agent_sdk.query() multi-turn through the bundled CLI, rather than the direct anthropic SDK. (Our original repro used the direct SDK to isolate the layer, but by construction that path never exercised the CLI, which is exactly where the bug lived — so it always looked healthy and couldn't confirm the fix.)

    Matrix, same endpoint / workspace / model (claude-sonnet-4-6), 3-turn runs with the real ~45K-token prefix (custom system prompt + full MCP tool array):

    SDK / bundled CLI Workaround cache_read (turn ≥2) Verdict
    0.1.81 / CLI 2.1.139 (our original broken version) none 77,533 ✅ healthy
    0.1.81 / CLI 2.1.139 CLAUDE_CODE_ATTRIBUTION_HEADER=0 110,412 ✅ healthy
    0.2.110 / CLI 2.1.191 none 90,295 ✅ healthy

    Per-turn it's the expected shape now: turn 1 writes the prefix, turn 2+ read it back cleanly (e.g. on 2.1.139: create 33,001 → read 33,001 → read 44,532). Previously this was cache_read=0 on every turn.

    Confirmed in production too: a real 50-turn multi-turn run read 1,883,932 tokens from cache with cache_create=109,750, stop_reason=end_turn, and no 429 — the rate-limit blow-through that this cache miss used to cause is gone.

    The CLAUDE_CODE_ATTRIBUTION_HEADER=0 workaround also works (and slightly reduces cache writes), but isn't necessary for us given the endpoint-side fix. Appreciate the root-cause detail on the per-request attestation token — matched our hypothesis #1 exactly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions