Skip to content

Make cost fields optional in ResultMessage to handle missing cost data - #3

Closed
r488it wants to merge 1 commit into
anthropics:mainfrom
r488it:feature/fix_cost_usd
Closed

r488it wants to merge 1 commit into
anthropics:mainfrom
r488it:feature/fix_cost_usd

Conversation

@r488it

@r488it r488it commented Jun 13, 2025

Copy link
Copy Markdown

Summary

Modified the ResultMessage class to make cost_usd and total_cost_usd fields optional, preventing errors when API responses don't include cost information.

Changes Made

  • Changed cost_usd field from float to float | None = None in ResultMessage class
  • Changed total_cost_usd field from float to float | None = None in ResultMessage class
  • Reordered fields to place optional fields at the end of the dataclass

Problem Solved

  • API responses sometimes don't include cost information (cost_usd, total_cost)
  • Current implementation has these fields commented out in client.py, which could cause initialization errors for ResultMessage
  • Need to handle messages reliably regardless of whether cost information is available

Files Modified

  • claude_code_sdk/types.py: Updated ResultMessage dataclass definition
  • claude_code_sdk/_internal/client.py: Maintained consistency with existing commented-out cost fields

Backward Compatibility

  • Existing code using cost information will need to add None checks
  • New implementation returns None when cost information is unavailable

Testing

  • Verified functionality with responses containing cost information
  • Verified functionality with responses missing cost information
  • Confirmed existing test cases still pass

Related Issue

This change improves the robustness of the SDK when handling different API response formats and prevents potential runtime errors due to missing cost data.

@musab-mk

Copy link
Copy Markdown

I got this on new Claude CLI update:

• SDK: Renamed total_cost to total_cost_usd


It might be related to that.

@106-

106- commented Jun 14, 2025

Copy link
Copy Markdown

As noted in the CHANGELOG.md, you can replace total_cost with total_cost_usd instead of just commenting it out.
I ran the actual command this library calls internally (tested with version 1.0.24), and got the following result:

$ claude --output-format stream-json --verbose --print "What is 2 + 2?"
...
{
  "type": "result",
  "subtype": "success",
  "is_error": false,
  "duration_ms": 2905,
  "duration_api_ms": 2702,
  "num_turns": 1,
  "result": "4",
  "session_id": "xxx",
  "total_cost_usd": 0.04995525000000001,
  "usage": {
    "input_tokens": 3,
    "cache_creation_input_tokens": 13291,
    "cache_read_input_tokens": 0,
    "output_tokens": 7,
    "server_tool_use": {
      "web_search_requests": 0
    }
  }
}

As you can see, total_cost_usd is indeed present in the response.

@ltawfik

ltawfik commented Jun 18, 2025

Copy link
Copy Markdown
Collaborator

thank you @r488it, this is fixed with the latest release, pip install --upgrade claude-code-sdk

@ltawfik ltawfik closed this Jun 18, 2025
golfballnut pushed a commit to golfballnut/claude-agent-sdk-python that referenced this pull request Oct 29, 2025
…cking, webhooks

Fixed 3 critical bugs causing 100% enrichment failure rate:

Bug anthropics#1: Contact Aggregation Logic
- Changed from REPLACE to MERGE strategy for fallback contacts
- Agent 2.1 (LinkedIn) and 2.2 (Perplexity) now accumulate contacts
- Lowered threshold from 2+ to 1+ contacts (proceed if ANY found)
- Only raises exception if 0 contacts from all sources
- Impact: Beach Buddy (1 contact) and Bear Lake (1 contact) now succeed

Bug anthropics#2: Missing course_id in Failures
- Added course_id to exception handler result dict
- Enables webhook firing even on enrichment failures
- Allows error tracking in ClickUp and database status updates
- Impact: All courses now tracked, even partial failures

Bug anthropics#3: Incorrect Cost Tracking
- Fixed Agent 1 cost reading from wrong variable
- Was: course_data.get("cost") (Agent 2's cost)
- Now: url_result.get("cost") (Agent 1's cost)
- Impact: Accurate budget tracking per agent

Tested against production logs showing 3 failed courses.
Expected: All should now complete with 1+ contacts.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
ollie-anthropic pushed a commit that referenced this pull request Dec 9, 2025
- Add SubagentExecutionConfig for controlling parallel vs sequential
  subagent execution when multiple Task tools are invoked in same turn
- Add execution_mode parameter to AgentDefinition for per-agent control
- Add subagent_execution field to ClaudeAgentOptions for global config
- Improve HookMatcher docstring with comprehensive examples including
  MCP tool matching patterns (mcp__server__tool format)
- Add new type aliases: SubagentExecutionMode, MultiInvocationMode,
  SubagentErrorHandling
- Add comprehensive tests for all new types

Addresses issues #2 (Critical: subagent execution mode) and #3 (Medium:
HookMatcher documentation).
bddppq referenced this pull request in bddppq/claude-agent-sdk-python Jun 9, 2026
win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that
together form a deterministic hook channel over the interactive CLI's own
settings.json command-hook mechanism.

- _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook
  event. Reads the hook-event JSON on stdin, forwards it (with an auth token)
  to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's
  hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no
  opinion") and exits 0, so a broken shim degrades to the CLI's normal
  behavior, never a hang/crash.
- _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback
  fallback) that dispatches each event to the user's options.hooks callbacks
  (matching the old control-protocol hook input/output shapes, incl. async_/
  continue_ conversion) and, for PreToolUse when can_use_tool is set, calls
  can_use_tool and translates PermissionResultAllow(updated_input=...)/
  PermissionResultDeny(message=...) into the CLI's
  hookSpecificOutput.permissionDecision + updatedInput short-circuit.
  build_hooks_settings() synthesizes the settings.json hooks block wiring the
  relevant events to the shim (preserving user matchers). Token-authenticated;
  bounded reads; all dispatch failures return `{}`.

Verified end-to-end (server + real shim subprocess): PreToolUse allow with
updated_input rewrites the input, deny carries the reason, PostToolUse returns
additionalContext, bad-token returns `{}`, user hooks fire with correct input
shapes + tool_use_id.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq referenced this pull request in bddppq/claude-agent-sdk-python Jun 9, 2026
win #3 (part 2/3): the transport now starts the HookIpcServer in connect(),
synthesizes a --settings hooks block wiring the relevant events to the shim,
injects the shim's IPC endpoint into the child env, and tears the server down
in close(). Programmatic hooks AND can_use_tool now run over this deterministic
channel instead of (only) the TUI screen-scraper.

pty_cli.py:
- connect(): _start_hook_ipc() before _build_command(); env injection of
  CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the
  synthesized hooks into an existing --settings JSON object (or appends one;
  non-object existing value -> second --settings, which the CLI merges).
- _validate_options(): removed the `hooks` rejection (now supported).
- _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the
  transport falls back to the existing watcher and hooks simply don't fire.
- close(): stop the hook IPC server (cancel serve task + unlink socket; no leak).
- Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission
  allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the
  watcher (the PreToolUse hook already decided server-side and the CLI
  short-circuits the dialog). _on_hook_permission_decision records a hook-channel
  deny into _pending_denied_tools so result.permission_denials still correlates.

_hook_ipc.py: added the optional on_permission_decision sink (fires after a
can_use_tool PreToolUse decision) so the transport can record hook-channel denies.

Tests:
- test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process
  dispatch (user hook shape, matcher filtering, async_/continue_ conversion,
  can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined
  user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip,
  bad-token and no-endpoint fail-open.
- test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI
  that actually runs the shim): PreToolUse fires with correct input shape +
  tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool
  deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool
  allow with updated_input ACTUALLY rewrites the executed input (content ->
  REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response.
- test_pty_integration validation test flipped: hooks accepted (not rejected);
  permission_prompt_tool_name still rejected.
- test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge)
  + TestCanUseToolViaHookGuard (no double-invoke; deny recorded).

Full suite 1220 passed / 5 skipped. ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq referenced this pull request in bddppq/claude-agent-sdk-python Jun 9, 2026
win #3 (part 3/3): update the PTY transport module docstring to reflect that
programmatic hooks and can_use_tool are now supported via the settings-hook IPC
bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher
is the fallback for when the bridge can't start (and always handles
plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks).

Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`:
- updated_input ACTUALLY applied: a Write the model issued with
  content="ORIGINAL_MODEL_CONTENT" was rewritten by
  can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"});
  the file on disk contained the REWRITTEN content (PostToolUse tool_response
  confirmed it). Baseline old SDK identical.
- can_use_tool deny via the hook channel: tool not run, full input + real toolu_
  id delivered, result.permission_denials carries the baseline shape.
- PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort,
  hook_event_name, permission_mode, session_id, tool_input, tool_name,
  tool_use_id, transcript_path]; PostToolUse fires with tool_response.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq referenced this pull request in bddppq/claude-agent-sdk-python Jun 9, 2026
…ment) (#1)

* Run the CLI in interactive mode over a PTY instead of headless pipes

Replace the default transport with a PTY-based one that drives the
*interactive* Claude Code CLI rather than the headless stream-json pipe
(formerly `-p`/`--print`). The CLI is launched attached to a pseudo-terminal,
prompts are typed in as real keystrokes, and responses are read by tailing the
session `.jsonl` transcript and translating its records back into the SDK
message dicts the parser already understands.

This is invisible to callers: the public `query()` / `ClaudeSDKClient` API is
unchanged.

Details:
- New `PtyCLITransport` (POSIX only): allocates a PTY in raw mode, spawns the
  interactive CLI in its own process group, drains terminal output, and tails
  `<config>/projects/<sanitized-cwd>/<session-id>.jsonl`.
- Reuses `SubprocessCLITransport`'s command/option logic, stripping the
  stream-json I/O flags that only apply to headless mode.
- Pre-seeds `hasCompletedOnboarding` and per-folder `hasTrustDialogAccepted`
  in the CLI config so the TUI drops straight to the prompt, and normalizes
  `IS_SANDBOX=1` when running as root.
- Acknowledges the initialize control request locally and synthesizes a
  `result` from the transcript's `turn_duration` record so existing
  one-shot and multi-turn flows terminate correctly. Interrupt types ESC.
- Surfaces an error result if the CLI exits early instead of ending silently.
- `SubprocessCLITransport` is retained for explicit/custom use.
- Adds unit tests for translation/command/env/keystroke logic, repoints
  existing transport mocks to the new default, and adds an example.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Make the interactive PTY transport the SDK's only transport

Fully replace the headless subprocess `-p`/stream-json transport with the
interactive PTY transport, and map every SDK control operation to its
interactive-TUI equivalent so the rest of the SDK keeps working.

Transport replacement:
- Extract the transport-agnostic logic (CLI discovery, option->flag command
  building, subprocess env + OTEL propagation, version check) into a shared
  `_cli_command` module. The command it builds is the interactive command:
  no stream-json I/O flags, no `--print`, and always `--session-id` so the
  transcript can be located.
- `PtyCLITransport` is now self-contained and the sole transport; delete
  `subprocess_cli.py`.

Control operations -> interactive equivalents (PtyCLITransport):
- interrupt -> ESC
- set_permission_mode -> shift+tab cycles default/acceptEdits/plan
- set_model -> `/model` slash command
- mcp_status / get_context_usage / rewind_files -> `/mcp` `/context` `/rewind`
- initialize and other control requests are acknowledged locally so the
  handshake proceeds.

Tests:
- Add `test_cli_command.py` covering command building, env (+OTEL, IS_SANDBOX),
  skills/sandbox/settings, discovery, and version check.
- Migrate transport mocks/patches to the new default; drop tests that were
  specific to the removed stream-json pipe (stdin streaming, stdout buffering,
  `--session-mirror` flag).
- Whole suite passes (943 passed, 3 skipped); ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add comprehensive drop-in tests for the interactive PTY transport

Prove the new transport behaves as a seamless replacement for the former
subprocess `-p` transport, exercised through the unchanged public API.

- test_pty_integration.py: end-to-end tests that spawn the real
  PtyCLITransport (PTY allocation, keystroke injection, transcript tailing,
  message translation) against a fake `claude` binary and drive it through
  `query()` and `ClaudeSDKClient`. Covers one-shot, multi-turn, tool-result
  surfacing, plain-prompt suppression, and early-exit error reporting. Runs
  under both asyncio and trio; no auth or real model needed.
- test_pty_transport.py: add deterministic unit tests for the transcript tail
  loop (translation, bookkeeping-record skipping, malformed lines, tool
  results), early-exit error result, and the control-request -> interactive
  mappings (permission-mode shift+tab cycling incl. wrap-around and the
  non-cyclable bypass case, set_model `/model`, `/mcp` `/context` `/rewind`),
  plus init/control-response ordering.
- pty_cli: extract the TUI warmup delay into a module-level `_WARMUP_SECONDS`
  so tests can shrink it; no behavior change in production.

Full suite: 972 passed, 3 skipped; ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Harden the interactive PTY transport against the review findings

Address the breakage cases surfaced by the nitpick review: fix what is fixable,
and convert the architecturally-unfixable cases (features that need the
bidirectional control channel the interactive CLI does not expose) from silent
no-ops / hangs into loud, actionable errors.

Fail loud instead of silently breaking (C1/C2/C3/C5, H9):
- _validate_options() rejects can_use_tool, hooks, in-process SDK MCP servers
  (external MCP still works), session_store, and permission_prompt_tool_name
  with a clear CLIConnectionError at connect(); include_partial_messages /
  include_hook_events / stderr warn rather than silently doing nothing.
- Unsupported control requests (mcp_status, get_context_usage, mcp_reconnect,
  mcp_toggle, stop_task, rewind_files, uncyclable permission modes) now return
  an error control_response instead of a fake success.

Correctness fixes:
- Prompt fidelity (H4/H5/H6): type prompts via bracketed paste so newlines and
  special characters are preserved; prepend a single space for a leading
  "/"/"!"/"#" so the TUI doesn't enter command mode; warn on dropped non-text
  (multimodal) blocks.
- Result fidelity (H1/H2/H3): synthesize the result statefully from the turn --
  final assistant text, summed usage, real num_turns, and is_error on
  refusal/error (total_cost_usd remains unavailable from the transcript).
- Concurrency (M1): serialize all PTY keystroke sequences with a write lock.
- Transcript discovery (L2/L3): resolve the path lazily with safe fallbacks
  (exact path -> our-session-id rglob -> newest in our project dir, gated on
  fork/resume and scoped so a concurrent unrelated session is never picked up).
- Compaction tolerance (L4): re-read on truncation, dedup records by uuid.
- Parser robustness (L1): coerce string/missing content and backfill model.
- Lifecycle (L6): atexit cleanup of child process groups + master fds; atomic
  ~/.claude.json update (temp file + os.replace) to avoid clobbering (H10).
- Carry parent_tool_use_id from the transcript (M9).

Documented limitations: POSIX-only (M7), stderr/max_buffer_size inert (M4/M5),
structured_output/total_cost_usd unavailable (M6), rewind-by-id unsupported (M8).

Tests: expand test_pty_transport.py (validation, result fidelity, dedup,
bracketed-paste fidelity, control errors) and test_pty_integration.py (prompt
round-trip with newline/leading-slash, fail-loud option). 997 passed, 3
skipped; ruff + mypy clean; one-shot and multi-turn verified live.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Correct the can_use_tool rejection rationale to the verified reason

I had described can_use_tool as failing because it "renders a TUI dialog the
SDK can't answer / hangs" -- that was unverified (this environment auto-approves
all tools, so a permission prompt can't be reproduced here).

The actual, code-verifiable reason: can_use_tool is invoked only when
transport.read_messages() yields a control_request with subtype can_use_tool
(query.py). The PTY transport never produces one, so the callback can never
fire regardless of whether the CLI prompts or auto-approves. Tools still run
in interactive mode; they just can't be gated through this callback.

Update the validation error and docstring accordingly. No behavior change.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add structured question detection/extraction for the PTY transport

The interactive TUI renders blocking questions (tool-permission prompts,
plan approval, AskUserQuestion forms, app-level confirmations) as on-screen
dialogs with no machine-readable signal. Add a detector that reconstructs the
rendered screen and parses the dialog into a structured DetectedQuestion.

- pty_question.py: dependency-free parser over reconstructed screen lines plus
  a pyte-backed screen reconstruction helper. Detection is gated on dialog
  chrome so numbered lists in ordinary model output are not false positives.
  Extracts kind, tool, target/command, question text, options (with action
  semantics), preview, and AskUserQuestion tab headers.
- pty_cli.py: maintain a pyte screen fed by the drain loop and expose
  detect_question(); degrades to None when pyte is absent.
- pyte is an optional dependency (the pty-introspect extra), mirroring the
  guarded-import pattern used for opentelemetry.
- Unit tests over golden frames for every question kind plus negatives.

Extraction was validated against the stream-json can_use_tool payload as
ground truth (tool name, target, command, options match) for Write, Bash,
Edit, ExitPlanMode, and AskUserQuestion.

* PTY: compute cost/usage/model_usage/stop_reason on result (C1, C2, H4, H5)

Reconstruct the stream-json ResultMessage data fields from the interactive
transcript so the PTY transport is a faithful drop-in:

- C2: dedup assistant usage by message.id (transcript writes multiple snapshots
  of the same message; naive summing over-counted cache tokens) and aggregate
  the distinct messages of the turn. New TurnUsageAccumulator in _usage.py.
- C1: populate total_cost_usd (computed client-side from per-message usage x a
  small model-pricing table), model_usage, stop_reason, and a non-None
  permission_denials list. total_cost_usd was silently None, crashing
  f"${msg.total_cost_usd:.4f}". Pricing validated against the documented
  stream-json ground truth (within ~0.3% of $0.18327675).
- H4: map assistant error/refusal to faithful result subtypes
  (error_max_turns, error_max_budget_usd, error_during_execution).
- H5: num_turns now uses the CLI's per-turn messageCount (the stream-json
  baseline counts API turns), not a cumulative result counter.
- Track the CLI's real session id from transcript records (groundwork for M2).

Adds dependency-free unit tests (test_pty_usage.py) and result-fidelity tests.
Updated two existing tests that asserted the pre-fix (buggy) num_turns
semantics to assert the corrected stream-json behavior.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: populate init/server_info and mcp_status; consistent session_id (C3, M1, M2, partial C4)

- C3/M1: build a full init payload from options (tools, mcp_servers, model,
  permissionMode, apiKeySource, slash_commands, output_style) and use it for
  both the system/init message and the `initialize` control response, so
  get_server_info() is no longer empty and the init message is no longer a
  5-field stub. The transcript carries no init record, so this is reconstructed
  from options + faithful defaults.
- partial C4: get_mcp_status now returns the configured servers with a
  `pending` status instead of erroring (no live state is observable over a PTY,
  so we report configured-but-unknown rather than fabricating "connected").
  The genuinely unobservable controls (get_context_usage, rewind_files,
  stop_task, mcp_toggle) still return a clear error.
- M2: track the CLI's real session id from transcript records; result and init
  use it. Per-message session_id already comes from each record's own sessionId.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: inherit cwd, pass user, restore entrypoint, surface user content (L2, L3, E1, M4)

- L2: pass cwd=None to Popen when the caller didn't set options.cwd so the child
  inherits the parent's working directory (the old transport's behavior),
  instead of pinning it to a Path.cwd() snapshot. Transcript-path computation
  still uses the concrete resolved cwd.
- L3: pass options.user to Popen so the child runs as the configured OS user,
  matching the old transport.
- E1: restore CLAUDE_CODE_ENTRYPOINT to "sdk-py" (was "sdk-py-pty") so telemetry
  is keyed identically to the stream-json baseline for drop-in consumers.
- M4: surface structured (list) user records even without a tool_result block
  (e.g. image/document content) instead of dropping them; only the plain-text
  echo of the typed prompt (string content) is suppressed.

Updated the two tests that asserted the pre-fix behavior.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: track live permission mode from transcript records (H1, L5)

The CLI writes a permission-mode transcript record whenever the live mode
changes. Read it as the source of truth for self._permission_mode, so:

- set_permission_mode computes shift+tab steps from the real current mode and is
  corrected automatically if the CLI's applied mode differs (H1 confirmation,
  L5 drift), without blocking the call on a transcript round-trip.
- set_model tracks the requested model id for follow-up calls.

permission-mode records are consumed for state only and not surfaced as SDK
messages. Adds a tracking test.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: answer tool-permission dialogs via the TUI detector (C5, C6)

The interactive CLI renders tool-permission and plan-approval prompts as
on-screen dialogs that block the turn. A background watcher polls the emulated
screen (pty_question), and on a blocking dialog:

- C5: routes the decision through can_use_tool when configured (mapping
  allow/deny to the matching dialog option), recording denials on the result's
  permission_denials list.
- C6: with no callback, answers "allow" so the turn completes instead of
  hanging forever on an unanswered prompt.

Answers are sent as the option digit + Enter; each dialog is answered at most
once (fingerprint dedup). AskUserQuestion / app dialogs need real content input
and are left to the consumer.

Consequently can_use_tool is no longer rejected at connect, and the
SDK-internal permission_prompt_tool_name="stdio" sentinel (set by the client
alongside can_use_tool) is accepted and not forwarded to the interactive CLI as
a flag it does not understand. A caller-supplied tool name is still rejected.

choose_option() (pure, unit-tested) picks the dialog option for a decision.
Adds question-parser and transport-level answering tests.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: honor options.max_buffer_size for the message buffer (M8)

max_buffer_size bounded a byte pipe in the old transport; here it bounds the
count of buffered message dicts (the closest interactive equivalent), instead
of a hardcoded 1000. Updated the validation note accordingly.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: parse structured_output from json_schema turns (H6)

When options.output_format is a json_schema format, the CLI constrains the
final assistant text to the schema, so parse that text into
result["structured_output"] (with a fenced-code-block recovery fallback),
mirroring the stream-json baseline. None when no schema is requested or the
text is not valid JSON. Unit-tested.

Also updates /tmp/dropin_todo.md with the full status of every item.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix PTY drop-in parity: num_turns, assistant-snapshot dedup, calibrated cost

C1/C2/H5 were producing wrong values vs live stream-json:

- H5 (num_turns): the turn_duration record's `messageCount` counts streamed
  assistant SNAPSHOTS, not API turns (PONG had messageCount=4 -> wrong
  num_turns=4). Replaced with `num_turns = (tool_result records) + 1`, derived
  and verified against 7 live stream-json data points. Live A/B now matches
  exactly: PONG 1==1, 2-write+2-read 5==5.

- C2 (dedup): added emit-side dedup of assistant messages by `message.id`
  (`_seen_assistant_ids`) so streaming snapshots that share an id but have
  distinct top-level uuids are not delivered to consumers multiple times.
  Usage was already deduped via TurnUsageAccumulator.

- C1 (cost): recalibrated cache multipliers to live ground truth. The nominal
  read=0.1 / 1h-write=2.0 over-estimated by ~1.4%; calibrated to read=0.10543,
  write=1.32145 (fit to 4 live single-call ResultMessage points so the read
  term cancels). Reproduces all 4 points within 0.06%.

- L1: tightened _WARMUP_SECONDS 3.0 -> 1.5 (live-verified the first prompt
  still submits reliably at 1.5s; even 1.0s worked across repeated runs).

Tests: updated test_pty_usage (calibrated rates + live-ground-truth assertion)
and test_pty_transport (num_turns = tool_results+1; assistant snapshot emitted
once). All gates pass (ruff, mypy, pytest: 1045 passed, 5 skipped).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix model_usage key, cost rates, interrupt, server_info (R1,R2,R4,R5,R6,R8,R9)

R1: emit the result usage breakdown under the camelCase wire key `modelUsage`
(message_parser reads data.get("modelUsage")) with camelCase per-model sub-keys
(inputTokens/outputTokens/cacheReadInputTokens/cacheCreationInputTokens/
webSearchRequests/costUSD/contextWindow/maxOutputTokens). Was snake_case ->
ResultMessage.model_usage always None. Live A/B: now populated, identical shape.

R2: revert cache multipliers to Anthropic's nominal rates (read 0.1, 5m-write
1.25, 1h-write 2.0). Reproduces the CLI per-model opus costUSD to 0.00% (was
+5.5% with the back-solved 0.10543/1.32145, which wrongly fit the result-level
total that bundles the unobservable haiku helper line).

R4: interrupt() now synthesizes a terminating result (error_during_execution,
is_error, result=None) after ESC so receive_response() ends instead of hanging.
Track per-turn start time and reset the result latch on each new prompt.

R5: add _build_server_info() returning the baseline initialize-control-response
key set (account/agents/available_output_styles/commands/models/output_style/
pid), separate from the system/init MESSAGE shape, so get_server_info() consumers
find commands/output_style/etc. pid + agents populated; catalogs empty-but-keyed.

R6: enrich system/init message: backfill resolved model from the first assistant
record; add agents/plugins/skills keys.

R8: honor --include-hook-events at the CLI level (was silently dropped). Verified
empirically the interactive CLI writes no hook records to the transcript, so hook
messages remain unsurfaceable (warning updated).

R9: bound _seen_uuids (OrderedDict LRU, cap 2048) to stop unbounded growth in
long multi-turn sessions.

Adds unit tests for each fix; live A/B verified R1/R2/R4/R5 against the old SDK.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix can_use_tool deny path (RR1,RR2,RR4,RR5)

RR1: a can_use_tool deny no longer hangs the turn. The interactive CLI
writes no turn_duration after a deny (it goes idle), so _emit_result never
fired and receive_response() deadlocked. _answer_question now marks the turn
deny-terminated and the tail loop synthesizes a terminating result
(subtype=success, is_error=False) once the rejected tool_result lands,
mirroring the baseline. Double-emit guarded in _emit_result/_emit_deny_result.

RR2: permission_denials now matches the baseline shape
{tool_name, tool_use_id, tool_input(full)}. tool_use blocks are indexed by id
while tailing; the rejected tool_result is correlated back to its tool_use to
recover the real tool_use_id and full original input (the TUI only exposed a
target string).

RR4: a fresh system/init now leads every turn (was emitted once at connect),
matching the baseline's per-turn ordering.

RR5: the rejected (permission-denied) tool_result is excluded from num_turns
(it is not a real model round-trip); genuine tool errors still count.

RR3: verified NOT a divergence -- live A/B shows the baseline also bypasses
can_use_tool for allowed_tools-preapproved tools (callback fires only on "ask"
decisions), which the PTY already matches.

All fixes live-verified against /tmp/venv_old. Unit tests added.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: surface all assistant tool_use/thinking blocks (RV1)

The interactive transcript writes each content block of one assistant
message as a SEPARATE record sharing one message.id (distinct uuids), and
same-id tool_use records are interleaved with the user/tool_result records
they trigger. The emit-side _seen_assistant_ids dedup suppressed every
same-id record after the first, dropping all tool_use (the agent's actions)
and thinking blocks from the consumer stream.

Live A/B vs the stream-json baseline shows it emits ONE AssistantMessage
per block-record (per-block granularity), NOT a merged message. Emit each
assistant block-record directly as its own AssistantMessage; keep usage
dedup by message.id and per-uuid true-duplicate dedup.

Before/after ToolUseBlock counts (3-file Write, baseline=3): PTY 0 -> 3.
Full message sequence and block counts now match the baseline exactly;
thinking blocks preserved. No regression to num_turns, cost/usage, the
deny path, per-turn init, or duplicate-message suppression.

Add TestDedup::test_multiblock_same_id_blocks_all_survive.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix empty-prompt hang, idle interrupt, resume context (RW1,RW2,RW3)

RW1: an empty/whitespace prompt cannot be submitted over the TUI, so no
turn ran and receive_response() hung forever. _handle_user_message now
synthesizes a terminating subtype=success result (mirroring the
deny/interrupt synthesis) via _emit_empty_prompt_result so the call
returns, matching the baseline subtype. Live A/B: query("")/query("   ")
-> subtype=success,is_error=False,num_turns=1 == baseline (was hang).

RW2: interrupt() while idle buffered a spurious error_during_execution
that the NEXT turn read first, corrupting it. _emit_interrupt_result now
no-ops when _turn_start_time is None (no active turn). Live A/B: idle
interrupt then query -> subtype=success,result='OK' == baseline (was
error_during_execution,result=None). Mid-turn interrupt (R4) unaffected.

RW3 (folds in M3): resume/continue did not restore context.
 - _cli_command no longer auto-appends --session-id over --resume/
   --continue unless the caller set one (matches baseline; the auto id
   made the CLI append to the resumed session while the transport tailed
   a nonexistent path).
 - _compute_transcript_path / new _resume_transcript_path resolve the
   actually-resumed file deterministically.
 - _tail_loop seeds offset from _initial_tail_offset (resume file size
   at connect) so the prior turn's turn_duration is not replayed as a
   stale result.
Live A/B: turn1 store 4271, resume in new transport, recall -> '4271'
== baseline (was stale 'DONE.').

Unit tests added for all three. No regressions (live A/B): RV1 tool_use
x3, num_turns (PONG 1, 3-file 4), usage 1897/5, per-turn init ordering,
deny path subtype=success/denials shape -- all match baseline.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add always-on pure-relay API monitor to enrich PTY transport output

Interpose a tiny loopback HTTP proxy between the interactive CLI and the
Anthropic API so the transport can observe the CLI<->API exchange the
.jsonl transcript cannot show. The proxy is ALWAYS ON (no config/env
toggle): connect() starts it on 127.0.0.1:<ephemeral>, captures the
original ANTHROPIC_BASE_URL as upstream, and points the child CLI's
ANTHROPIC_BASE_URL at the proxy; close() stops it.

PURE RELAY: request/response bytes are forwarded UNCHANGED (only the Host
header is rewritten loopback->upstream, mandatory framing — the API
rejects a private-IP Host with 403; body and all other headers are
byte-identical). SSE chunks are relayed as they arrive (never buffered),
so streaming to the CLI is preserved. Response framing (Content-Length /
chunked) is honored so a keep-alive connection cannot hang the relay. The
tee is non-fatal: forward first, then parse a gzip/chunk-decoded COPY of
/v1/messages behind try/except.

Use the captured traffic to enrich the synthesized result:
- C1/R3: total_cost_usd / model_usage / usage from the ACTUAL per-call
  usage across ALL /v1/messages calls in the turn, including the
  auxiliary helper-model call the transcript never records (no longer
  under-reporting the helper line).
- R7: duration_api_ms = sum of the real per-call request->response
  durations (no longer equal to the wall-clock duration_ms).
- C1: api_error_status surfaced when the turn ends on a non-2xx call
  (a retried-then-recovered transient is not surfaced).
- RV2: can_use_tool receives the FULL tool_use.input (and real
  tool_use_id) recovered from the intercepted response, correlated by
  tool name, instead of the scraped {target}.

The serve loop uses spawn_detached (not a held-open task group) so it is
safe to stop from any task. Falls back to transcript-derived usage and
the wall-clock duration when the monitor observes no traffic.

Unit tests: tests/test_api_monitor.py (relay transparency vs a fake
upstream, SSE tee, gzip, full tool_use input, non-/v1/messages not teed)
and TestApiMonitorEnrichment / TestRecoverToolInputRV2 in
tests/test_pty_transport.py (cost incl. helper call, summed
duration_api_ms, terminal-vs-recovered api_error_status, RV2 full input).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: fix RV2 race, quota over-count, keep-alive, exact endpoint (RL8/RL5/N1/N2/N3)

Iteration-2 review fixes for the always-on pure-relay API monitor:

RL8 / RV2 (deterministic): _decide_permission now bounded-awaits the relay
tee (_await_recovered_tool_input, <=2s/20ms poll) before falling back to the
scraped {target}. The CLI renders the permission dialog only after the full
/v1/messages response the relay also fully received, so the tee callback is
guaranteed within a small window. Live A/B: 5/5 deny runs deliver the full
{file_path,content} + real toolu_ id (was nondeterministic).

RL5 (cost reconciliation): root-caused the model_usage divergence. The
helper/title-gen call DOES traverse the relay; the interactive CLI just sends
it with the main model (opus) instead of haiku (small/fast-model resolution
differs from the headless baseline) -- a CLI-mode artifact, not a capture gap.
Per-call cost math is exact (haiku line matched baseline to the cent). Fixed
the real over-count: the synthetic quota_check probe (max_tokens:1, single
"quota" user message) is now relayed-but-not-teed so it can't inflate
cost/duration/error-status.

N1 (keep-alive): _relay_connection serves successive requests on one client
connection until close is signaled or framing is ambiguous, so a pooling
client never sees a mid-pool dropped connection.

N2 (exact endpoint): tee only when urlsplit(path).path == /v1/messages, so
/v1/messages/count_tokens and /batches are relayed-but-not-teed while
?beta=true still matches.

N3 (body overflow): _relay_request forwards exactly Content-Length body bytes
and carries the remainder as leftover for the next pipelined request.

Adds unit tests for RV2 await/correlate, exact endpoint match, quota-probe
exclusion, and keep-alive/pipelining. Full suite: 1109 passed, 5 skipped.
ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: traffic-derived enrichments RL9/RL10/RL11

RL9 (partial messages): when options.include_partial_messages is set, the
monitor tee reconstructs stream_event/StreamEvent messages from the SSE events
the relay already sees (message_start / content_block_* / message_delta /
message_stop), matching the stream-json baseline's StreamEvent shape. Emitted
via send_nowait from the monitor serve task (same event loop), best-effort and
non-blocking. Nothing is emitted when the option is unset (no behavior change).

RL10 (tools catalog): capture the CLI's real resolved tools list + model from
the teed /v1/messages request body; init/system message and get_server_info()
prefer this observed catalog over the thin options-derived defaults, so they
report what the model was actually offered.

RL11 (get_context_usage): the previously-unsupported control is now answered
from the latest successful per-call usage (input + cache_read + cache_creation
tokens => totalTokens), with percentage derived against the model context
window. Unobservable breakdowns are returned empty rather than fabricated.

Live A/B vs the stream-json baseline confirms each: same StreamEvent kinds,
real 29-tool catalog + model on get_server_info, real context token counts.
No regression: relay transparency intact, RV2 5/5 full deny input, cost /
num_turns / duration_api_ms unchanged. Unit tests added for all three.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: refresh get_server_info catalog, drop stale partial warning (RL10/RL9-warn/N4/N5)

RL10: client.get_server_info() now re-issues the `initialize` control request
(re-caching _initialization_result) and returns the fresh response instead of the
frozen connect-time snapshot, so the live tool catalog/model observed from
/v1/messages traffic are surfaced. Safe-if-not-connected; falls back to the cached
snapshot on refresh failure. Live: tools 0 (pre-turn) -> 29, model claude-opus-4-8.

RL9-warn: deleted the now-false include_partial_messages "has no effect" warning
in _validate_options (RL9 emits StreamEvents from the relay-teed SSE stream). The
adjacent include_hook_events/stderr warnings remain.

N4: documented the PTY/relay limitation in get_server_info()'s docstring (catalog
and model become available only after the first turn's traffic; connect-time init
is options-derived).

N5: kept the non-blocking send_nowait drop in _emit_stream_events (option a) and
documented why blocking would stall the relay serve task (transparency violation)
-- a necessary residual; buffer is generous and user-tunable.

Tests: get_server_info refresh/fallback/not-connected; include_partial_messages no
longer warns. Gates: ruff + mypy clean, pytest 1125 passed / 5 skipped.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: map updated_permissions to TUI session-allow (RL12)

When can_use_tool returns PermissionResultAllow with non-empty,
session-broad updated_permissions, the question-watcher presses the
dialog's allow_persist option ("Yes, allow all edits during this
session") so the persist intent survives the session, instead of the
plain allow_once option.

Over-grant guard (_should_persist_allow): press allow_persist only when
the dialog actually offers it AND the update is genuinely session-broad
(setMode, or an allow addRules/replaceRules scoped to session with no
narrowing rule_content). Narrow rules, disk destinations, deny/ask
behaviors, or empty updates fall back to allow_once -- never silently
grant broader-than-requested via the coarse TUI affordance.

Empty updated_permissions -> allow_once, unchanged.

Tests: TestPersistAllowRL12 (policy + _decide_permission +
_answer_question keystroke + dedup) and choose_option allow_persist
cases.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Tighten setMode over-grant guard in _should_persist_allow (RL13)

The setMode branch of _should_persist_allow returned True for ANY
setMode PermissionUpdate, so allowing a tool while requesting a
NARROWING mode (plan/default) would still press the TUI's "allow all
edits this session" option -- granting strictly broader than requested,
contradicting the guard's "never broader than requested" invariant.

Gate the setMode branch on a broadening target: only acceptEdits /
bypassPermissions map onto the session accept-edits press; plan /
default / None / other fall through to allow_once.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Harden _should_persist_allow into strict positive allowlist (RL14)

Close the empty/None-rules over-grant in the addRules/replaceRules branch
(rules=[] or None now `continue` instead of vacuously passing the
narrowing-rule check and pressing "allow all edits this session"), and
restructure the whole function into a strict positive allowlist: it
returns True ONLY for (a) setMode to a broadening mode
(acceptEdits/bypassPermissions), or (b) addRules/replaceRules with an
affirmative behavior=="allow" (None behavior now also falls through),
destination in (session, None), non-empty rules, and no narrowing
rule_content. Everything else continues to False.

Tests: empty-rules add/replaceRules (rules=[]/None x dest session/None)
-> allow_once; None behavior -> no persist; unknown update types -> no
persist. Existing RL12/RL13 cases re-confirmed.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Require unanimous session-broad list in _should_persist_allow (RL15)

Factor the single-element positive test into a static helper
_is_session_broad(upd) and stop short-circuiting in _should_persist_allow:
it now presses the TUI session-allow option only when EVERY element of
updated_permissions is itself session-broad (all(...)), so a mixed list
like [setMode acceptEdits, addRules deny] or [broad allow, narrow
rule_content] falls back to allow_once instead of dropping the narrowing
element. Unanimity is order-independent.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY transport: deliver child stderr via a dedicated pipe (H3)

The interactive child was spawned with stdin=stdout=stderr on the slave PTY,
so its stderr was muxed onto the terminal and the options.stderr callback never
fired (the old stream-json transport delivered stderr per line).

Give the child a SEPARATE stderr pipe when options.stderr is set: keep
stdin/stdout on the slave PTY (the TUI needs a tty) but route fd 2 to a pipe a
background reader splits into lines and hands to the callback, mirroring the old
transport's per-line behavior. The reader is non-fatal, and the fd + task are
cleaned up in close() (no leak, no orphan). With no callback, stderr stays on the
PTY (unchanged). Removed the now-false "stderr not invoked" warning from
_validate_options.

Live-verified: a real claude turn completes (stdout still on the PTY) and the
callback receives the child's real stderr line. Added an end-to-end integration
test (fake CLI writes to fd 2; turn still completes and lines are delivered).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY transport: isolate config writes + auto-accept bypass dialog (L4/anthropics#10)

L4 -- stop mutating the user's ~/.claude.json. The CLI persists OAuth/login
state and settings there and writes onboarding/trust flags into it, so a drop-in
must not churn or corrupt it. When the caller/env has NOT set CLAUDE_CONFIG_DIR,
default the child to a PERSISTENT SDK-owned dir ($XDG_CACHE_HOME or ~/.cache,
under claude-agent-sdk/config) so login survives across runs. Seed it once from
the user's real ~/.claude.json (best-effort copy) so auth + settings carry over
without ever writing the user's file; _ensure_onboarding_complete then writes the
flags into the SDK copy. The same effective config dir feeds transcript-path
resolution, so tailing still finds the session .jsonl. An explicit caller/env
CLAUDE_CONFIG_DIR is respected unchanged. options.env is never mutated in place.

anthropics#10 -- bypassPermissions startup dialog. The screen detector now flags the
"Bypass Permissions mode ... 1. No, exit / 2. Yes, I accept" dialog (is_bypass),
and the watcher auto-accepts it so the session does not hang. (As root the dialog
is suppressed by IS_SANDBOX=1; this covers the non-root case.) Other app dialogs
(e.g. folder-trust, handled by config pre-seeding) are left alone.

Live-verified: a turn completes against the isolated config dir (auth survived
the seeding -- real cost reported) and the user's ~/.claude.json is byte- and
mtime-identical before/after; a bypassPermissions turn completes without hanging.
Added unit tests for config-dir defaulting/seeding/user-file-untouched/respect,
bypass classification, and bypass auto-accept.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add settings-hook IPC bridge + shim for PTY transport

win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that
together form a deterministic hook channel over the interactive CLI's own
settings.json command-hook mechanism.

- _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook
  event. Reads the hook-event JSON on stdin, forwards it (with an auth token)
  to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's
  hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no
  opinion") and exits 0, so a broken shim degrades to the CLI's normal
  behavior, never a hang/crash.
- _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback
  fallback) that dispatches each event to the user's options.hooks callbacks
  (matching the old control-protocol hook input/output shapes, incl. async_/
  continue_ conversion) and, for PreToolUse when can_use_tool is set, calls
  can_use_tool and translates PermissionResultAllow(updated_input=...)/
  PermissionResultDeny(message=...) into the CLI's
  hookSpecificOutput.permissionDecision + updatedInput short-circuit.
  build_hooks_settings() synthesizes the settings.json hooks block wiring the
  relevant events to the shim (preserving user matchers). Token-authenticated;
  bounded reads; all dispatch failures return `{}`.

Verified end-to-end (server + real shim subprocess): PreToolUse allow with
updated_input rewrites the input, deny carries the reason, PostToolUse returns
additionalContext, bad-token returns `{}`, user hooks fire with correct input
shapes + tool_use_id.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Wire hooks + can_use_tool through the settings-hook IPC bridge (PTY)

win #3 (part 2/3): the transport now starts the HookIpcServer in connect(),
synthesizes a --settings hooks block wiring the relevant events to the shim,
injects the shim's IPC endpoint into the child env, and tears the server down
in close(). Programmatic hooks AND can_use_tool now run over this deterministic
channel instead of (only) the TUI screen-scraper.

pty_cli.py:
- connect(): _start_hook_ipc() before _build_command(); env injection of
  CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the
  synthesized hooks into an existing --settings JSON object (or appends one;
  non-object existing value -> second --settings, which the CLI merges).
- _validate_options(): removed the `hooks` rejection (now supported).
- _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the
  transport falls back to the existing watcher and hooks simply don't fire.
- close(): stop the hook IPC server (cancel serve task + unlink socket; no leak).
- Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission
  allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the
  watcher (the PreToolUse hook already decided server-side and the CLI
  short-circuits the dialog). _on_hook_permission_decision records a hook-channel
  deny into _pending_denied_tools so result.permission_denials still correlates.

_hook_ipc.py: added the optional on_permission_decision sink (fires after a
can_use_tool PreToolUse decision) so the transport can record hook-channel denies.

Tests:
- test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process
  dispatch (user hook shape, matcher filtering, async_/continue_ conversion,
  can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined
  user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip,
  bad-token and no-endpoint fail-open.
- test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI
  that actually runs the shim): PreToolUse fires with correct input shape +
  tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool
  deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool
  allow with updated_input ACTUALLY rewrites the executed input (content ->
  REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response.
- test_pty_integration validation test flipped: hooks accepted (not rejected);
  permission_prompt_tool_name still rejected.
- test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge)
  + TestCanUseToolViaHookGuard (no double-invoke; deny recorded).

Full suite 1220 passed / 5 skipped. ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Document hooks + can_use_tool support; live A/B verification

win #3 (part 3/3): update the PTY transport module docstring to reflect that
programmatic hooks and can_use_tool are now supported via the settings-hook IPC
bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher
is the fallback for when the bridge can't start (and always handles
plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks).

Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`:
- updated_input ACTUALLY applied: a Write the model issued with
  content="ORIGINAL_MODEL_CONTENT" was rewritten by
  can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"});
  the file on disk contained the REWRITTEN content (PostToolUse tool_response
  confirmed it). Baseline old SDK identical.
- can_use_tool deny via the hook channel: tool not run, full input + real toolu_
  id delivered, result.permission_denials carries the baseline shape.
- PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort,
  hook_event_name, permission_mode, session_id, tool_input, tool_name,
  tool_use_id, transcript_path]; PostToolUse fires with tool_response.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix hook-IPC review findings W1-W4 (PTY transport)

W1: gate the Unix-socket branch on hasattr(socket, "AF_UNIX") instead of the
always-False hasattr(os, "AF_UNIX"), so POSIX uses a 0700-dir Unix domain
socket (TCP loopback stays the fallback).

W2: close the can_use_tool permission bypass. build_hooks_settings now ALWAYS
emits a catch-all PreToolUse shim entry when can_use_tool is set (in addition to
any narrow user matchers), so the callback is consulted for every tool. The IPC
dispatch dedups can_use_tool per tool_use_id so a tool matching both the
catch-all and the user's narrow matcher invokes the callback only once.

W3: _close_stderr_pipe now closes both the write and read pipe ends (and clears
_stderr_read_fd) on every failed-spawn path, so a direct caller that skips
close() after a failed connect() does not leak the read fd.

W4: _matching_callbacks fires non-tool-event hooks (UserPromptSubmit/Stop/...)
regardless of the matcher pattern, mirroring the CLI's getMatchingHooks
(undefined matchQuery -> all matchers). Settings synthesis already omits the
tool matcher for non-tool events.

Tests: +catch-all/dedup/unmatched-tool/no-double-invoke (W2), unix-socket+0700
(W1), failed-spawn fd cleanup (W3), non-tool-event matcher fires (W4). Gates
green: ruff + mypy clean, pytest 1234 passed / 5 skipped.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix hook-IPC permission merge: deny-wins + per-id concurrency (W5/W6)

W5 (MEDIUM): a can_use_tool ALLOW could silently overwrite a user
PreToolUse hook DENY and leak a stale "hook-deny" reason into the allow
(self-contradictory, under-enforced). _dispatch now routes the decision
through HookIpcServer._merge_permission_decision: deny from EITHER source
wins, and the permission triplet (permissionDecision /
permissionDecisionReason / updatedInput) is replaced as an atomic unit so
no stale field leaks. Non-permission hookSpecificOutput fields are kept.
updated_input still applies on a both-allow result.

W6 (LOW-MED): _run_can_use_tool_deduped held _cut_lock across the user
can_use_tool callback, serializing distinct concurrent tool_use_ids (and
risking deadlock on re-entrancy). Switched to double-checked locking with
a per-tool_use_id _CutInflight marker: the lock now only guards the cache
/in-flight dicts, distinct ids run the callback concurrently, and a
duplicate fire for the same id awaits+replays the single result
(exactly-once consult). perm is pre-bound so a raising callback cannot
hang waiters. Bounded 512-entry FIFO cache retained.

Tests: W5 deny-wins matrix (reason correctness + updated_input-on-allow +
non-perm-field preservation); W6 distinct-id non-serialization,
same-id single-consult, and raising-callback-wakes-waiters.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

---------

Co-authored-by: Claude <noreply@anthropic.com>
bddppq referenced this pull request in bddppq/claude-agent-sdk-python Jun 9, 2026
…ment) (#1)

* Run the CLI in interactive mode over a PTY instead of headless pipes

Replace the default transport with a PTY-based one that drives the
*interactive* Claude Code CLI rather than the headless stream-json pipe
(formerly `-p`/`--print`). The CLI is launched attached to a pseudo-terminal,
prompts are typed in as real keystrokes, and responses are read by tailing the
session `.jsonl` transcript and translating its records back into the SDK
message dicts the parser already understands.

This is invisible to callers: the public `query()` / `ClaudeSDKClient` API is
unchanged.

Details:
- New `PtyCLITransport` (POSIX only): allocates a PTY in raw mode, spawns the
  interactive CLI in its own process group, drains terminal output, and tails
  `<config>/projects/<sanitized-cwd>/<session-id>.jsonl`.
- Reuses `SubprocessCLITransport`'s command/option logic, stripping the
  stream-json I/O flags that only apply to headless mode.
- Pre-seeds `hasCompletedOnboarding` and per-folder `hasTrustDialogAccepted`
  in the CLI config so the TUI drops straight to the prompt, and normalizes
  `IS_SANDBOX=1` when running as root.
- Acknowledges the initialize control request locally and synthesizes a
  `result` from the transcript's `turn_duration` record so existing
  one-shot and multi-turn flows terminate correctly. Interrupt types ESC.
- Surfaces an error result if the CLI exits early instead of ending silently.
- `SubprocessCLITransport` is retained for explicit/custom use.
- Adds unit tests for translation/command/env/keystroke logic, repoints
  existing transport mocks to the new default, and adds an example.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Make the interactive PTY transport the SDK's only transport

Fully replace the headless subprocess `-p`/stream-json transport with the
interactive PTY transport, and map every SDK control operation to its
interactive-TUI equivalent so the rest of the SDK keeps working.

Transport replacement:
- Extract the transport-agnostic logic (CLI discovery, option->flag command
  building, subprocess env + OTEL propagation, version check) into a shared
  `_cli_command` module. The command it builds is the interactive command:
  no stream-json I/O flags, no `--print`, and always `--session-id` so the
  transcript can be located.
- `PtyCLITransport` is now self-contained and the sole transport; delete
  `subprocess_cli.py`.

Control operations -> interactive equivalents (PtyCLITransport):
- interrupt -> ESC
- set_permission_mode -> shift+tab cycles default/acceptEdits/plan
- set_model -> `/model` slash command
- mcp_status / get_context_usage / rewind_files -> `/mcp` `/context` `/rewind`
- initialize and other control requests are acknowledged locally so the
  handshake proceeds.

Tests:
- Add `test_cli_command.py` covering command building, env (+OTEL, IS_SANDBOX),
  skills/sandbox/settings, discovery, and version check.
- Migrate transport mocks/patches to the new default; drop tests that were
  specific to the removed stream-json pipe (stdin streaming, stdout buffering,
  `--session-mirror` flag).
- Whole suite passes (943 passed, 3 skipped); ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add comprehensive drop-in tests for the interactive PTY transport

Prove the new transport behaves as a seamless replacement for the former
subprocess `-p` transport, exercised through the unchanged public API.

- test_pty_integration.py: end-to-end tests that spawn the real
  PtyCLITransport (PTY allocation, keystroke injection, transcript tailing,
  message translation) against a fake `claude` binary and drive it through
  `query()` and `ClaudeSDKClient`. Covers one-shot, multi-turn, tool-result
  surfacing, plain-prompt suppression, and early-exit error reporting. Runs
  under both asyncio and trio; no auth or real model needed.
- test_pty_transport.py: add deterministic unit tests for the transcript tail
  loop (translation, bookkeeping-record skipping, malformed lines, tool
  results), early-exit error result, and the control-request -> interactive
  mappings (permission-mode shift+tab cycling incl. wrap-around and the
  non-cyclable bypass case, set_model `/model`, `/mcp` `/context` `/rewind`),
  plus init/control-response ordering.
- pty_cli: extract the TUI warmup delay into a module-level `_WARMUP_SECONDS`
  so tests can shrink it; no behavior change in production.

Full suite: 972 passed, 3 skipped; ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Harden the interactive PTY transport against the review findings

Address the breakage cases surfaced by the nitpick review: fix what is fixable,
and convert the architecturally-unfixable cases (features that need the
bidirectional control channel the interactive CLI does not expose) from silent
no-ops / hangs into loud, actionable errors.

Fail loud instead of silently breaking (C1/C2/C3/C5, H9):
- _validate_options() rejects can_use_tool, hooks, in-process SDK MCP servers
  (external MCP still works), session_store, and permission_prompt_tool_name
  with a clear CLIConnectionError at connect(); include_partial_messages /
  include_hook_events / stderr warn rather than silently doing nothing.
- Unsupported control requests (mcp_status, get_context_usage, mcp_reconnect,
  mcp_toggle, stop_task, rewind_files, uncyclable permission modes) now return
  an error control_response instead of a fake success.

Correctness fixes:
- Prompt fidelity (H4/H5/H6): type prompts via bracketed paste so newlines and
  special characters are preserved; prepend a single space for a leading
  "/"/"!"/"#" so the TUI doesn't enter command mode; warn on dropped non-text
  (multimodal) blocks.
- Result fidelity (H1/H2/H3): synthesize the result statefully from the turn --
  final assistant text, summed usage, real num_turns, and is_error on
  refusal/error (total_cost_usd remains unavailable from the transcript).
- Concurrency (M1): serialize all PTY keystroke sequences with a write lock.
- Transcript discovery (L2/L3): resolve the path lazily with safe fallbacks
  (exact path -> our-session-id rglob -> newest in our project dir, gated on
  fork/resume and scoped so a concurrent unrelated session is never picked up).
- Compaction tolerance (L4): re-read on truncation, dedup records by uuid.
- Parser robustness (L1): coerce string/missing content and backfill model.
- Lifecycle (L6): atexit cleanup of child process groups + master fds; atomic
  ~/.claude.json update (temp file + os.replace) to avoid clobbering (H10).
- Carry parent_tool_use_id from the transcript (M9).

Documented limitations: POSIX-only (M7), stderr/max_buffer_size inert (M4/M5),
structured_output/total_cost_usd unavailable (M6), rewind-by-id unsupported (M8).

Tests: expand test_pty_transport.py (validation, result fidelity, dedup,
bracketed-paste fidelity, control errors) and test_pty_integration.py (prompt
round-trip with newline/leading-slash, fail-loud option). 997 passed, 3
skipped; ruff + mypy clean; one-shot and multi-turn verified live.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Correct the can_use_tool rejection rationale to the verified reason

I had described can_use_tool as failing because it "renders a TUI dialog the
SDK can't answer / hangs" -- that was unverified (this environment auto-approves
all tools, so a permission prompt can't be reproduced here).

The actual, code-verifiable reason: can_use_tool is invoked only when
transport.read_messages() yields a control_request with subtype can_use_tool
(query.py). The PTY transport never produces one, so the callback can never
fire regardless of whether the CLI prompts or auto-approves. Tools still run
in interactive mode; they just can't be gated through this callback.

Update the validation error and docstring accordingly. No behavior change.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add structured question detection/extraction for the PTY transport

The interactive TUI renders blocking questions (tool-permission prompts,
plan approval, AskUserQuestion forms, app-level confirmations) as on-screen
dialogs with no machine-readable signal. Add a detector that reconstructs the
rendered screen and parses the dialog into a structured DetectedQuestion.

- pty_question.py: dependency-free parser over reconstructed screen lines plus
  a pyte-backed screen reconstruction helper. Detection is gated on dialog
  chrome so numbered lists in ordinary model output are not false positives.
  Extracts kind, tool, target/command, question text, options (with action
  semantics), preview, and AskUserQuestion tab headers.
- pty_cli.py: maintain a pyte screen fed by the drain loop and expose
  detect_question(); degrades to None when pyte is absent.
- pyte is an optional dependency (the pty-introspect extra), mirroring the
  guarded-import pattern used for opentelemetry.
- Unit tests over golden frames for every question kind plus negatives.

Extraction was validated against the stream-json can_use_tool payload as
ground truth (tool name, target, command, options match) for Write, Bash,
Edit, ExitPlanMode, and AskUserQuestion.

* PTY: compute cost/usage/model_usage/stop_reason on result (C1, C2, H4, H5)

Reconstruct the stream-json ResultMessage data fields from the interactive
transcript so the PTY transport is a faithful drop-in:

- C2: dedup assistant usage by message.id (transcript writes multiple snapshots
  of the same message; naive summing over-counted cache tokens) and aggregate
  the distinct messages of the turn. New TurnUsageAccumulator in _usage.py.
- C1: populate total_cost_usd (computed client-side from per-message usage x a
  small model-pricing table), model_usage, stop_reason, and a non-None
  permission_denials list. total_cost_usd was silently None, crashing
  f"${msg.total_cost_usd:.4f}". Pricing validated against the documented
  stream-json ground truth (within ~0.3% of $0.18327675).
- H4: map assistant error/refusal to faithful result subtypes
  (error_max_turns, error_max_budget_usd, error_during_execution).
- H5: num_turns now uses the CLI's per-turn messageCount (the stream-json
  baseline counts API turns), not a cumulative result counter.
- Track the CLI's real session id from transcript records (groundwork for M2).

Adds dependency-free unit tests (test_pty_usage.py) and result-fidelity tests.
Updated two existing tests that asserted the pre-fix (buggy) num_turns
semantics to assert the corrected stream-json behavior.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: populate init/server_info and mcp_status; consistent session_id (C3, M1, M2, partial C4)

- C3/M1: build a full init payload from options (tools, mcp_servers, model,
  permissionMode, apiKeySource, slash_commands, output_style) and use it for
  both the system/init message and the `initialize` control response, so
  get_server_info() is no longer empty and the init message is no longer a
  5-field stub. The transcript carries no init record, so this is reconstructed
  from options + faithful defaults.
- partial C4: get_mcp_status now returns the configured servers with a
  `pending` status instead of erroring (no live state is observable over a PTY,
  so we report configured-but-unknown rather than fabricating "connected").
  The genuinely unobservable controls (get_context_usage, rewind_files,
  stop_task, mcp_toggle) still return a clear error.
- M2: track the CLI's real session id from transcript records; result and init
  use it. Per-message session_id already comes from each record's own sessionId.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: inherit cwd, pass user, restore entrypoint, surface user content (L2, L3, E1, M4)

- L2: pass cwd=None to Popen when the caller didn't set options.cwd so the child
  inherits the parent's working directory (the old transport's behavior),
  instead of pinning it to a Path.cwd() snapshot. Transcript-path computation
  still uses the concrete resolved cwd.
- L3: pass options.user to Popen so the child runs as the configured OS user,
  matching the old transport.
- E1: restore CLAUDE_CODE_ENTRYPOINT to "sdk-py" (was "sdk-py-pty") so telemetry
  is keyed identically to the stream-json baseline for drop-in consumers.
- M4: surface structured (list) user records even without a tool_result block
  (e.g. image/document content) instead of dropping them; only the plain-text
  echo of the typed prompt (string content) is suppressed.

Updated the two tests that asserted the pre-fix behavior.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: track live permission mode from transcript records (H1, L5)

The CLI writes a permission-mode transcript record whenever the live mode
changes. Read it as the source of truth for self._permission_mode, so:

- set_permission_mode computes shift+tab steps from the real current mode and is
  corrected automatically if the CLI's applied mode differs (H1 confirmation,
  L5 drift), without blocking the call on a transcript round-trip.
- set_model tracks the requested model id for follow-up calls.

permission-mode records are consumed for state only and not surfaced as SDK
messages. Adds a tracking test.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: answer tool-permission dialogs via the TUI detector (C5, C6)

The interactive CLI renders tool-permission and plan-approval prompts as
on-screen dialogs that block the turn. A background watcher polls the emulated
screen (pty_question), and on a blocking dialog:

- C5: routes the decision through can_use_tool when configured (mapping
  allow/deny to the matching dialog option), recording denials on the result's
  permission_denials list.
- C6: with no callback, answers "allow" so the turn completes instead of
  hanging forever on an unanswered prompt.

Answers are sent as the option digit + Enter; each dialog is answered at most
once (fingerprint dedup). AskUserQuestion / app dialogs need real content input
and are left to the consumer.

Consequently can_use_tool is no longer rejected at connect, and the
SDK-internal permission_prompt_tool_name="stdio" sentinel (set by the client
alongside can_use_tool) is accepted and not forwarded to the interactive CLI as
a flag it does not understand. A caller-supplied tool name is still rejected.

choose_option() (pure, unit-tested) picks the dialog option for a decision.
Adds question-parser and transport-level answering tests.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: honor options.max_buffer_size for the message buffer (M8)

max_buffer_size bounded a byte pipe in the old transport; here it bounds the
count of buffered message dicts (the closest interactive equivalent), instead
of a hardcoded 1000. Updated the validation note accordingly.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY: parse structured_output from json_schema turns (H6)

When options.output_format is a json_schema format, the CLI constrains the
final assistant text to the schema, so parse that text into
result["structured_output"] (with a fenced-code-block recovery fallback),
mirroring the stream-json baseline. None when no schema is requested or the
text is not valid JSON. Unit-tested.

Also updates /tmp/dropin_todo.md with the full status of every item.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix PTY drop-in parity: num_turns, assistant-snapshot dedup, calibrated cost

C1/C2/H5 were producing wrong values vs live stream-json:

- H5 (num_turns): the turn_duration record's `messageCount` counts streamed
  assistant SNAPSHOTS, not API turns (PONG had messageCount=4 -> wrong
  num_turns=4). Replaced with `num_turns = (tool_result records) + 1`, derived
  and verified against 7 live stream-json data points. Live A/B now matches
  exactly: PONG 1==1, 2-write+2-read 5==5.

- C2 (dedup): added emit-side dedup of assistant messages by `message.id`
  (`_seen_assistant_ids`) so streaming snapshots that share an id but have
  distinct top-level uuids are not delivered to consumers multiple times.
  Usage was already deduped via TurnUsageAccumulator.

- C1 (cost): recalibrated cache multipliers to live ground truth. The nominal
  read=0.1 / 1h-write=2.0 over-estimated by ~1.4%; calibrated to read=0.10543,
  write=1.32145 (fit to 4 live single-call ResultMessage points so the read
  term cancels). Reproduces all 4 points within 0.06%.

- L1: tightened _WARMUP_SECONDS 3.0 -> 1.5 (live-verified the first prompt
  still submits reliably at 1.5s; even 1.0s worked across repeated runs).

Tests: updated test_pty_usage (calibrated rates + live-ground-truth assertion)
and test_pty_transport (num_turns = tool_results+1; assistant snapshot emitted
once). All gates pass (ruff, mypy, pytest: 1045 passed, 5 skipped).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix model_usage key, cost rates, interrupt, server_info (R1,R2,R4,R5,R6,R8,R9)

R1: emit the result usage breakdown under the camelCase wire key `modelUsage`
(message_parser reads data.get("modelUsage")) with camelCase per-model sub-keys
(inputTokens/outputTokens/cacheReadInputTokens/cacheCreationInputTokens/
webSearchRequests/costUSD/contextWindow/maxOutputTokens). Was snake_case ->
ResultMessage.model_usage always None. Live A/B: now populated, identical shape.

R2: revert cache multipliers to Anthropic's nominal rates (read 0.1, 5m-write
1.25, 1h-write 2.0). Reproduces the CLI per-model opus costUSD to 0.00% (was
+5.5% with the back-solved 0.10543/1.32145, which wrongly fit the result-level
total that bundles the unobservable haiku helper line).

R4: interrupt() now synthesizes a terminating result (error_during_execution,
is_error, result=None) after ESC so receive_response() ends instead of hanging.
Track per-turn start time and reset the result latch on each new prompt.

R5: add _build_server_info() returning the baseline initialize-control-response
key set (account/agents/available_output_styles/commands/models/output_style/
pid), separate from the system/init MESSAGE shape, so get_server_info() consumers
find commands/output_style/etc. pid + agents populated; catalogs empty-but-keyed.

R6: enrich system/init message: backfill resolved model from the first assistant
record; add agents/plugins/skills keys.

R8: honor --include-hook-events at the CLI level (was silently dropped). Verified
empirically the interactive CLI writes no hook records to the transcript, so hook
messages remain unsurfaceable (warning updated).

R9: bound _seen_uuids (OrderedDict LRU, cap 2048) to stop unbounded growth in
long multi-turn sessions.

Adds unit tests for each fix; live A/B verified R1/R2/R4/R5 against the old SDK.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix can_use_tool deny path (RR1,RR2,RR4,RR5)

RR1: a can_use_tool deny no longer hangs the turn. The interactive CLI
writes no turn_duration after a deny (it goes idle), so _emit_result never
fired and receive_response() deadlocked. _answer_question now marks the turn
deny-terminated and the tail loop synthesizes a terminating result
(subtype=success, is_error=False) once the rejected tool_result lands,
mirroring the baseline. Double-emit guarded in _emit_result/_emit_deny_result.

RR2: permission_denials now matches the baseline shape
{tool_name, tool_use_id, tool_input(full)}. tool_use blocks are indexed by id
while tailing; the rejected tool_result is correlated back to its tool_use to
recover the real tool_use_id and full original input (the TUI only exposed a
target string).

RR4: a fresh system/init now leads every turn (was emitted once at connect),
matching the baseline's per-turn ordering.

RR5: the rejected (permission-denied) tool_result is excluded from num_turns
(it is not a real model round-trip); genuine tool errors still count.

RR3: verified NOT a divergence -- live A/B shows the baseline also bypasses
can_use_tool for allowed_tools-preapproved tools (callback fires only on "ask"
decisions), which the PTY already matches.

All fixes live-verified against /tmp/venv_old. Unit tests added.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: surface all assistant tool_use/thinking blocks (RV1)

The interactive transcript writes each content block of one assistant
message as a SEPARATE record sharing one message.id (distinct uuids), and
same-id tool_use records are interleaved with the user/tool_result records
they trigger. The emit-side _seen_assistant_ids dedup suppressed every
same-id record after the first, dropping all tool_use (the agent's actions)
and thinking blocks from the consumer stream.

Live A/B vs the stream-json baseline shows it emits ONE AssistantMessage
per block-record (per-block granularity), NOT a merged message. Emit each
assistant block-record directly as its own AssistantMessage; keep usage
dedup by message.id and per-uuid true-duplicate dedup.

Before/after ToolUseBlock counts (3-file Write, baseline=3): PTY 0 -> 3.
Full message sequence and block counts now match the baseline exactly;
thinking blocks preserved. No regression to num_turns, cost/usage, the
deny path, per-turn init, or duplicate-message suppression.

Add TestDedup::test_multiblock_same_id_blocks_all_survive.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY drop-in: fix empty-prompt hang, idle interrupt, resume context (RW1,RW2,RW3)

RW1: an empty/whitespace prompt cannot be submitted over the TUI, so no
turn ran and receive_response() hung forever. _handle_user_message now
synthesizes a terminating subtype=success result (mirroring the
deny/interrupt synthesis) via _emit_empty_prompt_result so the call
returns, matching the baseline subtype. Live A/B: query("")/query("   ")
-> subtype=success,is_error=False,num_turns=1 == baseline (was hang).

RW2: interrupt() while idle buffered a spurious error_during_execution
that the NEXT turn read first, corrupting it. _emit_interrupt_result now
no-ops when _turn_start_time is None (no active turn). Live A/B: idle
interrupt then query -> subtype=success,result='OK' == baseline (was
error_during_execution,result=None). Mid-turn interrupt (R4) unaffected.

RW3 (folds in M3): resume/continue did not restore context.
 - _cli_command no longer auto-appends --session-id over --resume/
   --continue unless the caller set one (matches baseline; the auto id
   made the CLI append to the resumed session while the transport tailed
   a nonexistent path).
 - _compute_transcript_path / new _resume_transcript_path resolve the
   actually-resumed file deterministically.
 - _tail_loop seeds offset from _initial_tail_offset (resume file size
   at connect) so the prior turn's turn_duration is not replayed as a
   stale result.
Live A/B: turn1 store 4271, resume in new transport, recall -> '4271'
== baseline (was stale 'DONE.').

Unit tests added for all three. No regressions (live A/B): RV1 tool_use
x3, num_turns (PONG 1, 3-file 4), usage 1897/5, per-turn init ordering,
deny path subtype=success/denials shape -- all match baseline.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add always-on pure-relay API monitor to enrich PTY transport output

Interpose a tiny loopback HTTP proxy between the interactive CLI and the
Anthropic API so the transport can observe the CLI<->API exchange the
.jsonl transcript cannot show. The proxy is ALWAYS ON (no config/env
toggle): connect() starts it on 127.0.0.1:<ephemeral>, captures the
original ANTHROPIC_BASE_URL as upstream, and points the child CLI's
ANTHROPIC_BASE_URL at the proxy; close() stops it.

PURE RELAY: request/response bytes are forwarded UNCHANGED (only the Host
header is rewritten loopback->upstream, mandatory framing — the API
rejects a private-IP Host with 403; body and all other headers are
byte-identical). SSE chunks are relayed as they arrive (never buffered),
so streaming to the CLI is preserved. Response framing (Content-Length /
chunked) is honored so a keep-alive connection cannot hang the relay. The
tee is non-fatal: forward first, then parse a gzip/chunk-decoded COPY of
/v1/messages behind try/except.

Use the captured traffic to enrich the synthesized result:
- C1/R3: total_cost_usd / model_usage / usage from the ACTUAL per-call
  usage across ALL /v1/messages calls in the turn, including the
  auxiliary helper-model call the transcript never records (no longer
  under-reporting the helper line).
- R7: duration_api_ms = sum of the real per-call request->response
  durations (no longer equal to the wall-clock duration_ms).
- C1: api_error_status surfaced when the turn ends on a non-2xx call
  (a retried-then-recovered transient is not surfaced).
- RV2: can_use_tool receives the FULL tool_use.input (and real
  tool_use_id) recovered from the intercepted response, correlated by
  tool name, instead of the scraped {target}.

The serve loop uses spawn_detached (not a held-open task group) so it is
safe to stop from any task. Falls back to transcript-derived usage and
the wall-clock duration when the monitor observes no traffic.

Unit tests: tests/test_api_monitor.py (relay transparency vs a fake
upstream, SSE tee, gzip, full tool_use input, non-/v1/messages not teed)
and TestApiMonitorEnrichment / TestRecoverToolInputRV2 in
tests/test_pty_transport.py (cost incl. helper call, summed
duration_api_ms, terminal-vs-recovered api_error_status, RV2 full input).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: fix RV2 race, quota over-count, keep-alive, exact endpoint (RL8/RL5/N1/N2/N3)

Iteration-2 review fixes for the always-on pure-relay API monitor:

RL8 / RV2 (deterministic): _decide_permission now bounded-awaits the relay
tee (_await_recovered_tool_input, <=2s/20ms poll) before falling back to the
scraped {target}. The CLI renders the permission dialog only after the full
/v1/messages response the relay also fully received, so the tee callback is
guaranteed within a small window. Live A/B: 5/5 deny runs deliver the full
{file_path,content} + real toolu_ id (was nondeterministic).

RL5 (cost reconciliation): root-caused the model_usage divergence. The
helper/title-gen call DOES traverse the relay; the interactive CLI just sends
it with the main model (opus) instead of haiku (small/fast-model resolution
differs from the headless baseline) -- a CLI-mode artifact, not a capture gap.
Per-call cost math is exact (haiku line matched baseline to the cent). Fixed
the real over-count: the synthetic quota_check probe (max_tokens:1, single
"quota" user message) is now relayed-but-not-teed so it can't inflate
cost/duration/error-status.

N1 (keep-alive): _relay_connection serves successive requests on one client
connection until close is signaled or framing is ambiguous, so a pooling
client never sees a mid-pool dropped connection.

N2 (exact endpoint): tee only when urlsplit(path).path == /v1/messages, so
/v1/messages/count_tokens and /batches are relayed-but-not-teed while
?beta=true still matches.

N3 (body overflow): _relay_request forwards exactly Content-Length body bytes
and carries the remainder as leftover for the next pipelined request.

Adds unit tests for RV2 await/correlate, exact endpoint match, quota-probe
exclusion, and keep-alive/pipelining. Full suite: 1109 passed, 5 skipped.
ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: traffic-derived enrichments RL9/RL10/RL11

RL9 (partial messages): when options.include_partial_messages is set, the
monitor tee reconstructs stream_event/StreamEvent messages from the SSE events
the relay already sees (message_start / content_block_* / message_delta /
message_stop), matching the stream-json baseline's StreamEvent shape. Emitted
via send_nowait from the monitor serve task (same event loop), best-effort and
non-blocking. Nothing is emitted when the option is unset (no behavior change).

RL10 (tools catalog): capture the CLI's real resolved tools list + model from
the teed /v1/messages request body; init/system message and get_server_info()
prefer this observed catalog over the thin options-derived defaults, so they
report what the model was actually offered.

RL11 (get_context_usage): the previously-unsupported control is now answered
from the latest successful per-call usage (input + cache_read + cache_creation
tokens => totalTokens), with percentage derived against the model context
window. Unobservable breakdowns are returned empty rather than fabricated.

Live A/B vs the stream-json baseline confirms each: same StreamEvent kinds,
real 29-tool catalog + model on get_server_info, real context token counts.
No regression: relay transparency intact, RV2 5/5 full deny input, cost /
num_turns / duration_api_ms unchanged. Unit tests added for all three.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: refresh get_server_info catalog, drop stale partial warning (RL10/RL9-warn/N4/N5)

RL10: client.get_server_info() now re-issues the `initialize` control request
(re-caching _initialization_result) and returns the fresh response instead of the
frozen connect-time snapshot, so the live tool catalog/model observed from
/v1/messages traffic are surfaced. Safe-if-not-connected; falls back to the cached
snapshot on refresh failure. Live: tools 0 (pre-turn) -> 29, model claude-opus-4-8.

RL9-warn: deleted the now-false include_partial_messages "has no effect" warning
in _validate_options (RL9 emits StreamEvents from the relay-teed SSE stream). The
adjacent include_hook_events/stderr warnings remain.

N4: documented the PTY/relay limitation in get_server_info()'s docstring (catalog
and model become available only after the first turn's traffic; connect-time init
is options-derived).

N5: kept the non-blocking send_nowait drop in _emit_stream_events (option a) and
documented why blocking would stall the relay serve task (transparency violation)
-- a necessary residual; buffer is generous and user-tunable.

Tests: get_server_info refresh/fallback/not-connected; include_partial_messages no
longer warns. Gates: ruff + mypy clean, pytest 1125 passed / 5 skipped.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY relay: map updated_permissions to TUI session-allow (RL12)

When can_use_tool returns PermissionResultAllow with non-empty,
session-broad updated_permissions, the question-watcher presses the
dialog's allow_persist option ("Yes, allow all edits during this
session") so the persist intent survives the session, instead of the
plain allow_once option.

Over-grant guard (_should_persist_allow): press allow_persist only when
the dialog actually offers it AND the update is genuinely session-broad
(setMode, or an allow addRules/replaceRules scoped to session with no
narrowing rule_content). Narrow rules, disk destinations, deny/ask
behaviors, or empty updates fall back to allow_once -- never silently
grant broader-than-requested via the coarse TUI affordance.

Empty updated_permissions -> allow_once, unchanged.

Tests: TestPersistAllowRL12 (policy + _decide_permission +
_answer_question keystroke + dedup) and choose_option allow_persist
cases.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Tighten setMode over-grant guard in _should_persist_allow (RL13)

The setMode branch of _should_persist_allow returned True for ANY
setMode PermissionUpdate, so allowing a tool while requesting a
NARROWING mode (plan/default) would still press the TUI's "allow all
edits this session" option -- granting strictly broader than requested,
contradicting the guard's "never broader than requested" invariant.

Gate the setMode branch on a broadening target: only acceptEdits /
bypassPermissions map onto the session accept-edits press; plan /
default / None / other fall through to allow_once.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Harden _should_persist_allow into strict positive allowlist (RL14)

Close the empty/None-rules over-grant in the addRules/replaceRules branch
(rules=[] or None now `continue` instead of vacuously passing the
narrowing-rule check and pressing "allow all edits this session"), and
restructure the whole function into a strict positive allowlist: it
returns True ONLY for (a) setMode to a broadening mode
(acceptEdits/bypassPermissions), or (b) addRules/replaceRules with an
affirmative behavior=="allow" (None behavior now also falls through),
destination in (session, None), non-empty rules, and no narrowing
rule_content. Everything else continues to False.

Tests: empty-rules add/replaceRules (rules=[]/None x dest session/None)
-> allow_once; None behavior -> no persist; unknown update types -> no
persist. Existing RL12/RL13 cases re-confirmed.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Require unanimous session-broad list in _should_persist_allow (RL15)

Factor the single-element positive test into a static helper
_is_session_broad(upd) and stop short-circuiting in _should_persist_allow:
it now presses the TUI session-allow option only when EVERY element of
updated_permissions is itself session-broad (all(...)), so a mixed list
like [setMode acceptEdits, addRules deny] or [broad allow, narrow
rule_content] falls back to allow_once instead of dropping the narrowing
element. Unanimity is order-independent.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY transport: deliver child stderr via a dedicated pipe (H3)

The interactive child was spawned with stdin=stdout=stderr on the slave PTY,
so its stderr was muxed onto the terminal and the options.stderr callback never
fired (the old stream-json transport delivered stderr per line).

Give the child a SEPARATE stderr pipe when options.stderr is set: keep
stdin/stdout on the slave PTY (the TUI needs a tty) but route fd 2 to a pipe a
background reader splits into lines and hands to the callback, mirroring the old
transport's per-line behavior. The reader is non-fatal, and the fd + task are
cleaned up in close() (no leak, no orphan). With no callback, stderr stays on the
PTY (unchanged). Removed the now-false "stderr not invoked" warning from
_validate_options.

Live-verified: a real claude turn completes (stdout still on the PTY) and the
callback receives the child's real stderr line. Added an end-to-end integration
test (fake CLI writes to fd 2; turn still completes and lines are delivered).

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* PTY transport: isolate config writes + auto-accept bypass dialog (L4/anthropics#10)

L4 -- stop mutating the user's ~/.claude.json. The CLI persists OAuth/login
state and settings there and writes onboarding/trust flags into it, so a drop-in
must not churn or corrupt it. When the caller/env has NOT set CLAUDE_CONFIG_DIR,
default the child to a PERSISTENT SDK-owned dir ($XDG_CACHE_HOME or ~/.cache,
under claude-agent-sdk/config) so login survives across runs. Seed it once from
the user's real ~/.claude.json (best-effort copy) so auth + settings carry over
without ever writing the user's file; _ensure_onboarding_complete then writes the
flags into the SDK copy. The same effective config dir feeds transcript-path
resolution, so tailing still finds the session .jsonl. An explicit caller/env
CLAUDE_CONFIG_DIR is respected unchanged. options.env is never mutated in place.

anthropics#10 -- bypassPermissions startup dialog. The screen detector now flags the
"Bypass Permissions mode ... 1. No, exit / 2. Yes, I accept" dialog (is_bypass),
and the watcher auto-accepts it so the session does not hang. (As root the dialog
is suppressed by IS_SANDBOX=1; this covers the non-root case.) Other app dialogs
(e.g. folder-trust, handled by config pre-seeding) are left alone.

Live-verified: a turn completes against the isolated config dir (auth survived
the seeding -- real cost reported) and the user's ~/.claude.json is byte- and
mtime-identical before/after; a bypassPermissions turn completes without hanging.
Added unit tests for config-dir defaulting/seeding/user-file-untouched/respect,
bypass classification, and bypass auto-accept.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Add settings-hook IPC bridge + shim for PTY transport

win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that
together form a deterministic hook channel over the interactive CLI's own
settings.json command-hook mechanism.

- _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook
  event. Reads the hook-event JSON on stdin, forwards it (with an auth token)
  to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's
  hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no
  opinion") and exits 0, so a broken shim degrades to the CLI's normal
  behavior, never a hang/crash.
- _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback
  fallback) that dispatches each event to the user's options.hooks callbacks
  (matching the old control-protocol hook input/output shapes, incl. async_/
  continue_ conversion) and, for PreToolUse when can_use_tool is set, calls
  can_use_tool and translates PermissionResultAllow(updated_input=...)/
  PermissionResultDeny(message=...) into the CLI's
  hookSpecificOutput.permissionDecision + updatedInput short-circuit.
  build_hooks_settings() synthesizes the settings.json hooks block wiring the
  relevant events to the shim (preserving user matchers). Token-authenticated;
  bounded reads; all dispatch failures return `{}`.

Verified end-to-end (server + real shim subprocess): PreToolUse allow with
updated_input rewrites the input, deny carries the reason, PostToolUse returns
additionalContext, bad-token returns `{}`, user hooks fire with correct input
shapes + tool_use_id.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Wire hooks + can_use_tool through the settings-hook IPC bridge (PTY)

win #3 (part 2/3): the transport now starts the HookIpcServer in connect(),
synthesizes a --settings hooks block wiring the relevant events to the shim,
injects the shim's IPC endpoint into the child env, and tears the server down
in close(). Programmatic hooks AND can_use_tool now run over this deterministic
channel instead of (only) the TUI screen-scraper.

pty_cli.py:
- connect(): _start_hook_ipc() before _build_command(); env injection of
  CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the
  synthesized hooks into an existing --settings JSON object (or appends one;
  non-object existing value -> second --settings, which the CLI merges).
- _validate_options(): removed the `hooks` rejection (now supported).
- _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the
  transport falls back to the existing watcher and hooks simply don't fire.
- close(): stop the hook IPC server (cancel serve task + unlink socket; no leak).
- Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission
  allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the
  watcher (the PreToolUse hook already decided server-side and the CLI
  short-circuits the dialog). _on_hook_permission_decision records a hook-channel
  deny into _pending_denied_tools so result.permission_denials still correlates.

_hook_ipc.py: added the optional on_permission_decision sink (fires after a
can_use_tool PreToolUse decision) so the transport can record hook-channel denies.

Tests:
- test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process
  dispatch (user hook shape, matcher filtering, async_/continue_ conversion,
  can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined
  user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip,
  bad-token and no-endpoint fail-open.
- test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI
  that actually runs the shim): PreToolUse fires with correct input shape +
  tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool
  deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool
  allow with updated_input ACTUALLY rewrites the executed input (content ->
  REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response.
- test_pty_integration validation test flipped: hooks accepted (not rejected);
  permission_prompt_tool_name still rejected.
- test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge)
  + TestCanUseToolViaHookGuard (no double-invoke; deny recorded).

Full suite 1220 passed / 5 skipped. ruff + mypy clean.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Document hooks + can_use_tool support; live A/B verification

win #3 (part 3/3): update the PTY transport module docstring to reflect that
programmatic hooks and can_use_tool are now supported via the settings-hook IPC
bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher
is the fallback for when the bridge can't start (and always handles
plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks).

Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`:
- updated_input ACTUALLY applied: a Write the model issued with
  content="ORIGINAL_MODEL_CONTENT" was rewritten by
  can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"});
  the file on disk contained the REWRITTEN content (PostToolUse tool_response
  confirmed it). Baseline old SDK identical.
- can_use_tool deny via the hook channel: tool not run, full input + real toolu_
  id delivered, result.permission_denials carries the baseline shape.
- PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort,
  hook_event_name, permission_mode, session_id, tool_input, tool_name,
  tool_use_id, transcript_path]; PostToolUse fires with tool_response.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix hook-IPC review findings W1-W4 (PTY transport)

W1: gate the Unix-socket branch on hasattr(socket, "AF_UNIX") instead of the
always-False hasattr(os, "AF_UNIX"), so POSIX uses a 0700-dir Unix domain
socket (TCP loopback stays the fallback).

W2: close the can_use_tool permission bypass. build_hooks_settings now ALWAYS
emits a catch-all PreToolUse shim entry when can_use_tool is set (in addition to
any narrow user matchers), so the callback is consulted for every tool. The IPC
dispatch dedups can_use_tool per tool_use_id so a tool matching both the
catch-all and the user's narrow matcher invokes the callback only once.

W3: _close_stderr_pipe now closes both the write and read pipe ends (and clears
_stderr_read_fd) on every failed-spawn path, so a direct caller that skips
close() after a failed connect() does not leak the read fd.

W4: _matching_callbacks fires non-tool-event hooks (UserPromptSubmit/Stop/...)
regardless of the matcher pattern, mirroring the CLI's getMatchingHooks
(undefined matchQuery -> all matchers). Settings synthesis already omits the
tool matcher for non-tool events.

Tests: +catch-all/dedup/unmatched-tool/no-double-invoke (W2), unix-socket+0700
(W1), failed-spawn fd cleanup (W3), non-tool-event matcher fires (W4). Gates
green: ruff + mypy clean, pytest 1234 passed / 5 skipped.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

* Fix hook-IPC permission merge: deny-wins + per-id concurrency (W5/W6)

W5 (MEDIUM): a can_use_tool ALLOW could silently overwrite a user
PreToolUse hook DENY and leak a stale "hook-deny" reason into the allow
(self-contradictory, under-enforced). _dispatch now routes the decision
through HookIpcServer._merge_permission_decision: deny from EITHER source
wins, and the permission triplet (permissionDecision /
permissionDecisionReason / updatedInput) is replaced as an atomic unit so
no stale field leaks. Non-permission hookSpecificOutput fields are kept.
updated_input still applies on a both-allow result.

W6 (LOW-MED): _run_can_use_tool_deduped held _cut_lock across the user
can_use_tool callback, serializing distinct concurrent tool_use_ids (and
risking deadlock on re-entrancy). Switched to double-checked locking with
a per-tool_use_id _CutInflight marker: the lock now only guards the cache
/in-flight dicts, distinct ids run the callback concurrently, and a
duplicate fire for the same id awaits+replays the single result
(exactly-once consult). perm is pre-bound so a raising callback cannot
hang waiters. Bounded 512-entry FIFO cache retained.

Tests: W5 deny-wins matrix (reason correctness + updated_input-on-allow +
non-perm-field preservation); W6 distinct-id non-serialization,
same-id single-consult, and raising-callback-wakes-waiters.

https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants