Repository navigation
Bundled CLI never reads prompt cache in multi-turn headless runs (direct anthropic SDK over the same endpoint caches correctly) #974
Description
Activity
Confirmando este bug en producción independientemente.
Observamos el mismo patrón en headless multi-turn con
claude-agent-sdk==0.2.82(Sonnet 4.6, CLI 2.1.142) lanzados como Server tmux con--dangerously-skip-permissions --print:- Cada turno escribe ~14-15K tokens
cache_creation_input_tokensy lee 0 tokenscache_read_input_tokens. - Aprox. a los 2-3 turnos disparamos 429 por token-rate.
- Misma key + mismo modelo + misma session usados con
anthropicSDK directo (messages.createconcache_control: {type: "ephemeral"}explícito) sí registracache_read_input_tokens > 0en los turnos 2+.
Esto coincide con tu repro al 100%: el caching layer funciona en el endpoint, lo que no lo activa correctamente es el bundled CLI cuando lo invoca el SDK.
Aprovecho para preguntar: ¿hay un ETA público para fix, o un workaround oficial recomendado mientras tanto (e.g. flag, env var, o downgrade a una versión SDK/CLI donde sí cacheaba)? Para nosotros decide entre seguir migrando a
claude-agent-sdko quedarnos enanthropicSDK directo con tool-loop manual.Gracias por el repro tan claro — facilita mucho debug del lado del consumidor.
- Cada turno escribe ~14-15K tokens
Thanks for the independent confirmation, @alexisoller-cpu - useful to have this off the Claude-on-AWS endpoint, since it rules out gateway-specific behaviour as the cause. The numerical fingerprint matches exactly on our side too.
It might be worth mentioning that pinning back to
claude-agent-sdk==0.1.81(bundled CLI 2.1.139) had the samecache_read=0. So a simple SDK/CLI downgrade is not the workaround.Our current mitigations are:
- Gating heavy MCP servers per-task so most runs only load the essential tools (smaller
toolsarray → smaller per-turn write). - Trimming the tool allowlist down to only what we actually call, in case the SDK filters at the API-request level (still verifying).
- Requesting an ITPM cap raise on our workspace so multi-turn runs complete despite the cache miss.
- Gating heavy MCP servers per-task so most runs only load the essential tools (smaller
Quick update with new evidence.
After @alexisoller-cpu's independent confirmation (same symptom on direct
Anthropic, not just on our Claude-on-AWS endpoint), I re-ran the probe with
the Cache Diagnostics beta
(cache-diagnosis-2026-04-07) attached and threaded
diagnostics.previous_message_idthrough the two calls.Result on the direct-SDK probe (same endpoint, same workspace, same model
as the broken worker run):[call 1] usage = {... cache_creation_input_tokens=2722, cache_read_input_tokens=0 ...} [call 1] diagnostics = None [call 2] usage = {... cache_creation_input_tokens=4, cache_read_input_tokens=2722 ...} [call 2] diagnostics = NonePer the Cache Diagnostics docs,
diagnostics: nullon call 2 (with a real
previous_message_id) means "a comparison ran and found no divergence."
So we now have two independent positive signals on the direct-SDK side:- The cache layer actually delivered cached tokens (
cache_read=2722). - Cache Diagnostics formally confirms the request fingerprint is byte-stable
across turns (diagnostics=null, nocache_miss_reason).
Together those eliminate "gateway / workspace / model / API caching / our
own request stability" as possible explanations. The only thing left that
differs between the healthy direct-SDK case and the broken
claude-agent-sdk → bundled CLIcase is whatever the CLI is doing to its
outbound request payload between turns.(Side note for anyone else who lands here: Cache Diagnostics works over the
https://aws-external-anthropic.<region>.api.awsendpoint — the doc only
excludes Bedrock and Vertex, and we confirmed it's honoured on Claude-on-AWS.)Concrete ask for SDK-side triage: the official next step the Anthropic
docs themselves point at for this class of bug is to enable
cache-diagnosis-2026-04-07and passdiagnostics.previous_message_id
between turns. If the bundled CLI's request loop did that on multi-turn
runs (or if there's an env var / option to surface it), the response's
cache_miss_reason.typefield would directly name the divergent component —
model_changed/system_changed/tools_changed/messages_changed—
and the fix follows from there. The hypotheses in the original issue
(per-turn injection into the system block, tool re-ordering after async MCP
load, breakpoint placement drift) all map cleanly onto one of those four
types, so the diagnostic immediately distinguishes them.Happy to share the probe script or any worker-side logs.
- The cache layer actually delivered cached tokens (
Thanks for the detailed report. You should be able to unblock yourself immediately by adding this to your options.env:
env={ ..., "CLAUDE_CODE_ATTRIBUTION_HEADER": "0" }This would drop a block which invalidates the prompt cache block. We are preparing the fix-forward for an upcoming SDK release so this will no longer be necessary on future releases.
Verified fixed as of bundled CLI 2.1.181 (re-tested on
claude-agent-sdk0.2.110 / CLI 2.1.195): across a multi-turn runcache_read_input_tokensis now non-zero and grows as the prefix accumulates whilecache_creationshrinks — prefix caching is working.Root cause matches the CLI 2.1.181 CHANGELOG entry — "Fixed prompt caching not reading on custom
ANTHROPIC_BASE_URLand on Foundry due to a per-request attestation token changing every turn" — which is exactly the per-turn dynamic prefix change hypothesized in this issue:
https://github.com/anthropics/claude-code/blob/01f1617f14452ac78bf319cef2236d87c0fe05cb/CHANGELOG.md#L159Upgrade
claude-agent-sdkto a release bundling CLI ≥ 2.1.181 to pick up the fix.@qing-ant, looks like the fix for this issue has been verified by @Magic-Man-us. We can close this bug.
Confirming the fix on our side — many thanks.
TL;DR: prefix caching is working again over the Claude-on-AWS first-party endpoint (
aws-external-anthropic.<region>.api.aws+anthropic-workspace-idheader), and — notably — it's healthy even on the same bundled CLI that originally reproduced the bug, without upgrading. That suggests the fix propagated server/gateway-side for this endpoint, not only via the CLI ≥ 2.1.181 bump.To verify I wrote a small CLI-path two-turn probe — driving
claude_agent_sdk.query()multi-turn through the bundled CLI, rather than the directanthropicSDK. (Our original repro used the direct SDK to isolate the layer, but by construction that path never exercised the CLI, which is exactly where the bug lived — so it always looked healthy and couldn't confirm the fix.)Matrix, same endpoint / workspace / model (
claude-sonnet-4-6), 3-turn runs with the real ~45K-token prefix (custom system prompt + full MCP tool array):SDK / bundled CLI Workaround cache_read(turn ≥2)Verdict 0.1.81 / CLI 2.1.139 (our original broken version) none 77,533 ✅ healthy 0.1.81 / CLI 2.1.139 CLAUDE_CODE_ATTRIBUTION_HEADER=0110,412 ✅ healthy 0.2.110 / CLI 2.1.191 none 90,295 ✅ healthy Per-turn it's the expected shape now: turn 1 writes the prefix, turn 2+ read it back cleanly (e.g. on 2.1.139: create 33,001 → read 33,001 → read 44,532). Previously this was
cache_read=0on every turn.Confirmed in production too: a real 50-turn multi-turn run read 1,883,932 tokens from cache with
cache_create=109,750,stop_reason=end_turn, and no 429 — the rate-limit blow-through that this cache miss used to cause is gone.The
CLAUDE_CODE_ATTRIBUTION_HEADER=0workaround also works (and slightly reduces cache writes), but isn't necessary for us given the endpoint-side fix. Appreciate the root-cause detail on the per-request attestation token — matched our hypothesis #1 exactly.
Summary
In a headless
claude-agent-sdkrun (Python, custom system prompt, MCPservers configured via
.mcp.json), the bundled Claude Code CLI writesprompt cache on every turn but never reads it —
cache_read_input_tokens=0across multi-turn investigations. On the same endpoint, same model, same
workspace credentials, a 30-line script using the plain
anthropicPythonSDK with
cache_control: ephemeralon the system block gets a healthycache read on the second call.
This makes the regression CLI-side rather than gateway- or API-side, and
the cache_create-on-every-turn behaviour quickly bursts past per-minute
input-token rate limits on multi-turn runs.
Environment
claude-agent-sdkversions tested: 0.1.81 (bundled CLI 2.1.139) and0.2.82 (bundled CLI 2.1.142) — both broken.
the bundled CLI shebang).
claude-sonnet-4-6.(
https://aws-external-anthropic.<region>.api.aws) with ananthropic-workspace-idcustom header. Same endpoint used by the healthydirect-SDK probe below.
.mcp.json(two HTTP-transport servers).permission_mode="bypassPermissions",strict_mcp_config=True, customsystem_promptstring (not theclaude_codepreset),cwdscoped to aread-only mount,
allowed_tools/disallowed_toolsenforced natively.Symptom — bundled CLI (broken)
Representative log line from a 3-turn run (annotated):
cache_create_input_tokens≈ 14–15K per turn, accumulating to43–46K across 3 turns.
cache_read_input_tokens= 0 across every turn.CLI exits with
is_error=true, subtype=success, api_error_status=429.The "ResultMessage…subtype=success api_error_status=429" combo is
reported via PR #923 — that part is working correctly. The cache miss
is the underlying issue.
Symptom — direct
anthropicSDK over the same endpoint (healthy)Same
ANTHROPIC_BASE_URL, sameanthropic-workspace-idheader, sameclaude-sonnet-4-6model, same ~11K-char system block carryingcache_control: {"type": "ephemeral"}.Conclusion: the gateway / API / workspace / model are all fine. The
bundled CLI is shipping a payload that defeats the cache lookup despite
the prefix appearing identical across turns.
Minimal repro
Two pieces:
1. Healthy reference — direct SDK call, ~50 LOC
2. Broken reference — same workspace, same endpoint, via
claude-agent-sdkThe model takes a couple of cheap tool calls and the second/third turn's
usage.cache_read_input_tokensis still 0.What we ruled out
(bundled CLI 2.1.139); same broken behaviour.
v0.1.81 exhibits the same symptom.
anthropicSDK over the sameendpoint shows healthy cache reads.
MCP_CONNECTION_NONBLOCKING=0— env var was set; no observableeffect (we may have had the env var name wrong, but it didn't move the
needle either way).
Hypotheses
Things the CLI may be doing differently from the direct SDK call, any of
which would defeat the prefix match:
token, tool registration metadata, or similar non-stable text appearing
inside the
cache_control-marked block.settle.
cache_controlbreakpoint placement drifting between turn 1 andturn 2 (e.g., placed on different block indices once more conversation
accumulates), so it doesn't match what was originally written.
subprocess CLI rejects list-form system_prompt) —if the CLI is forced to flatten
systemto a string, thecache_controlmarker placement may not survive intact.We can't observe which is true from outside the bundled binary. If anyone
on the SDK side can confirm what the CLI sends as
systemandtoolsbetween turns when given a custom
system_promptplus MCP servers, thatwould settle it quickly.
Impact
input-token cap will trip 429 within 2–3 turns even on small (<20K-token)
prefixes.
amortised cost.
defaults are lower than direct Anthropic accounts.
Asks
system block in headless mode, and if so, document the env var or
option to suppress it (analogous to
exclude_dynamic_sectionson thepreset).
be sufficient to diagnose from outside.
Happy to share the worker-side and probe scripts in their entirety if it
helps shorten triage.