Repository navigation
Conversation
|
I got this on new Claude CLI update: • SDK: Renamed It might be related to that. |
|
As noted in the CHANGELOG.md, you can replace As you can see, |
Collaborator
|
thank you @r488it, this is fixed with the latest release, |
golfballnut
pushed a commit
to golfballnut/claude-agent-sdk-python
that referenced
this pull request
Oct 29, 2025
…cking, webhooks Fixed 3 critical bugs causing 100% enrichment failure rate: Bug anthropics#1: Contact Aggregation Logic - Changed from REPLACE to MERGE strategy for fallback contacts - Agent 2.1 (LinkedIn) and 2.2 (Perplexity) now accumulate contacts - Lowered threshold from 2+ to 1+ contacts (proceed if ANY found) - Only raises exception if 0 contacts from all sources - Impact: Beach Buddy (1 contact) and Bear Lake (1 contact) now succeed Bug anthropics#2: Missing course_id in Failures - Added course_id to exception handler result dict - Enables webhook firing even on enrichment failures - Allows error tracking in ClickUp and database status updates - Impact: All courses now tracked, even partial failures Bug anthropics#3: Incorrect Cost Tracking - Fixed Agent 1 cost reading from wrong variable - Was: course_data.get("cost") (Agent 2's cost) - Now: url_result.get("cost") (Agent 1's cost) - Impact: Accurate budget tracking per agent Tested against production logs showing 3 failed courses. Expected: All should now complete with 1+ contacts. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
ollie-anthropic
pushed a commit
that referenced
this pull request
Dec 9, 2025
- Add SubagentExecutionConfig for controlling parallel vs sequential subagent execution when multiple Task tools are invoked in same turn - Add execution_mode parameter to AgentDefinition for per-agent control - Add subagent_execution field to ClaudeAgentOptions for global config - Improve HookMatcher docstring with comprehensive examples including MCP tool matching patterns (mcp__server__tool format) - Add new type aliases: SubagentExecutionMode, MultiInvocationMode, SubagentErrorHandling - Add comprehensive tests for all new types Addresses issues #2 (Critical: subagent execution mode) and #3 (Medium: HookMatcher documentation).
bddppq
referenced
this pull request
in bddppq/claude-agent-sdk-python
Jun 9, 2026
win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that together form a deterministic hook channel over the interactive CLI's own settings.json command-hook mechanism. - _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook event. Reads the hook-event JSON on stdin, forwards it (with an auth token) to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no opinion") and exits 0, so a broken shim degrades to the CLI's normal behavior, never a hang/crash. - _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback fallback) that dispatches each event to the user's options.hooks callbacks (matching the old control-protocol hook input/output shapes, incl. async_/ continue_ conversion) and, for PreToolUse when can_use_tool is set, calls can_use_tool and translates PermissionResultAllow(updated_input=...)/ PermissionResultDeny(message=...) into the CLI's hookSpecificOutput.permissionDecision + updatedInput short-circuit. build_hooks_settings() synthesizes the settings.json hooks block wiring the relevant events to the shim (preserving user matchers). Token-authenticated; bounded reads; all dispatch failures return `{}`. Verified end-to-end (server + real shim subprocess): PreToolUse allow with updated_input rewrites the input, deny carries the reason, PostToolUse returns additionalContext, bad-token returns `{}`, user hooks fire with correct input shapes + tool_use_id. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq
referenced
this pull request
in bddppq/claude-agent-sdk-python
Jun 9, 2026
win #3 (part 2/3): the transport now starts the HookIpcServer in connect(), synthesizes a --settings hooks block wiring the relevant events to the shim, injects the shim's IPC endpoint into the child env, and tears the server down in close(). Programmatic hooks AND can_use_tool now run over this deterministic channel instead of (only) the TUI screen-scraper. pty_cli.py: - connect(): _start_hook_ipc() before _build_command(); env injection of CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the synthesized hooks into an existing --settings JSON object (or appends one; non-object existing value -> second --settings, which the CLI merges). - _validate_options(): removed the `hooks` rejection (now supported). - _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the transport falls back to the existing watcher and hooks simply don't fire. - close(): stop the hook IPC server (cancel serve task + unlink socket; no leak). - Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the watcher (the PreToolUse hook already decided server-side and the CLI short-circuits the dialog). _on_hook_permission_decision records a hook-channel deny into _pending_denied_tools so result.permission_denials still correlates. _hook_ipc.py: added the optional on_permission_decision sink (fires after a can_use_tool PreToolUse decision) so the transport can record hook-channel denies. Tests: - test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process dispatch (user hook shape, matcher filtering, async_/continue_ conversion, can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip, bad-token and no-endpoint fail-open. - test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI that actually runs the shim): PreToolUse fires with correct input shape + tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool allow with updated_input ACTUALLY rewrites the executed input (content -> REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response. - test_pty_integration validation test flipped: hooks accepted (not rejected); permission_prompt_tool_name still rejected. - test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge) + TestCanUseToolViaHookGuard (no double-invoke; deny recorded). Full suite 1220 passed / 5 skipped. ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq
referenced
this pull request
in bddppq/claude-agent-sdk-python
Jun 9, 2026
win #3 (part 3/3): update the PTY transport module docstring to reflect that programmatic hooks and can_use_tool are now supported via the settings-hook IPC bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher is the fallback for when the bridge can't start (and always handles plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks). Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`: - updated_input ACTUALLY applied: a Write the model issued with content="ORIGINAL_MODEL_CONTENT" was rewritten by can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"}); the file on disk contained the REWRITTEN content (PostToolUse tool_response confirmed it). Baseline old SDK identical. - can_use_tool deny via the hook channel: tool not run, full input + real toolu_ id delivered, result.permission_denials carries the baseline shape. - PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort, hook_event_name, permission_mode, session_id, tool_input, tool_name, tool_use_id, transcript_path]; PostToolUse fires with tool_response. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH
bddppq
referenced
this pull request
in bddppq/claude-agent-sdk-python
Jun 9, 2026
…ment) (#1) * Run the CLI in interactive mode over a PTY instead of headless pipes Replace the default transport with a PTY-based one that drives the *interactive* Claude Code CLI rather than the headless stream-json pipe (formerly `-p`/`--print`). The CLI is launched attached to a pseudo-terminal, prompts are typed in as real keystrokes, and responses are read by tailing the session `.jsonl` transcript and translating its records back into the SDK message dicts the parser already understands. This is invisible to callers: the public `query()` / `ClaudeSDKClient` API is unchanged. Details: - New `PtyCLITransport` (POSIX only): allocates a PTY in raw mode, spawns the interactive CLI in its own process group, drains terminal output, and tails `<config>/projects/<sanitized-cwd>/<session-id>.jsonl`. - Reuses `SubprocessCLITransport`'s command/option logic, stripping the stream-json I/O flags that only apply to headless mode. - Pre-seeds `hasCompletedOnboarding` and per-folder `hasTrustDialogAccepted` in the CLI config so the TUI drops straight to the prompt, and normalizes `IS_SANDBOX=1` when running as root. - Acknowledges the initialize control request locally and synthesizes a `result` from the transcript's `turn_duration` record so existing one-shot and multi-turn flows terminate correctly. Interrupt types ESC. - Surfaces an error result if the CLI exits early instead of ending silently. - `SubprocessCLITransport` is retained for explicit/custom use. - Adds unit tests for translation/command/env/keystroke logic, repoints existing transport mocks to the new default, and adds an example. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Make the interactive PTY transport the SDK's only transport Fully replace the headless subprocess `-p`/stream-json transport with the interactive PTY transport, and map every SDK control operation to its interactive-TUI equivalent so the rest of the SDK keeps working. Transport replacement: - Extract the transport-agnostic logic (CLI discovery, option->flag command building, subprocess env + OTEL propagation, version check) into a shared `_cli_command` module. The command it builds is the interactive command: no stream-json I/O flags, no `--print`, and always `--session-id` so the transcript can be located. - `PtyCLITransport` is now self-contained and the sole transport; delete `subprocess_cli.py`. Control operations -> interactive equivalents (PtyCLITransport): - interrupt -> ESC - set_permission_mode -> shift+tab cycles default/acceptEdits/plan - set_model -> `/model` slash command - mcp_status / get_context_usage / rewind_files -> `/mcp` `/context` `/rewind` - initialize and other control requests are acknowledged locally so the handshake proceeds. Tests: - Add `test_cli_command.py` covering command building, env (+OTEL, IS_SANDBOX), skills/sandbox/settings, discovery, and version check. - Migrate transport mocks/patches to the new default; drop tests that were specific to the removed stream-json pipe (stdin streaming, stdout buffering, `--session-mirror` flag). - Whole suite passes (943 passed, 3 skipped); ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add comprehensive drop-in tests for the interactive PTY transport Prove the new transport behaves as a seamless replacement for the former subprocess `-p` transport, exercised through the unchanged public API. - test_pty_integration.py: end-to-end tests that spawn the real PtyCLITransport (PTY allocation, keystroke injection, transcript tailing, message translation) against a fake `claude` binary and drive it through `query()` and `ClaudeSDKClient`. Covers one-shot, multi-turn, tool-result surfacing, plain-prompt suppression, and early-exit error reporting. Runs under both asyncio and trio; no auth or real model needed. - test_pty_transport.py: add deterministic unit tests for the transcript tail loop (translation, bookkeeping-record skipping, malformed lines, tool results), early-exit error result, and the control-request -> interactive mappings (permission-mode shift+tab cycling incl. wrap-around and the non-cyclable bypass case, set_model `/model`, `/mcp` `/context` `/rewind`), plus init/control-response ordering. - pty_cli: extract the TUI warmup delay into a module-level `_WARMUP_SECONDS` so tests can shrink it; no behavior change in production. Full suite: 972 passed, 3 skipped; ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Harden the interactive PTY transport against the review findings Address the breakage cases surfaced by the nitpick review: fix what is fixable, and convert the architecturally-unfixable cases (features that need the bidirectional control channel the interactive CLI does not expose) from silent no-ops / hangs into loud, actionable errors. Fail loud instead of silently breaking (C1/C2/C3/C5, H9): - _validate_options() rejects can_use_tool, hooks, in-process SDK MCP servers (external MCP still works), session_store, and permission_prompt_tool_name with a clear CLIConnectionError at connect(); include_partial_messages / include_hook_events / stderr warn rather than silently doing nothing. - Unsupported control requests (mcp_status, get_context_usage, mcp_reconnect, mcp_toggle, stop_task, rewind_files, uncyclable permission modes) now return an error control_response instead of a fake success. Correctness fixes: - Prompt fidelity (H4/H5/H6): type prompts via bracketed paste so newlines and special characters are preserved; prepend a single space for a leading "/"/"!"/"#" so the TUI doesn't enter command mode; warn on dropped non-text (multimodal) blocks. - Result fidelity (H1/H2/H3): synthesize the result statefully from the turn -- final assistant text, summed usage, real num_turns, and is_error on refusal/error (total_cost_usd remains unavailable from the transcript). - Concurrency (M1): serialize all PTY keystroke sequences with a write lock. - Transcript discovery (L2/L3): resolve the path lazily with safe fallbacks (exact path -> our-session-id rglob -> newest in our project dir, gated on fork/resume and scoped so a concurrent unrelated session is never picked up). - Compaction tolerance (L4): re-read on truncation, dedup records by uuid. - Parser robustness (L1): coerce string/missing content and backfill model. - Lifecycle (L6): atexit cleanup of child process groups + master fds; atomic ~/.claude.json update (temp file + os.replace) to avoid clobbering (H10). - Carry parent_tool_use_id from the transcript (M9). Documented limitations: POSIX-only (M7), stderr/max_buffer_size inert (M4/M5), structured_output/total_cost_usd unavailable (M6), rewind-by-id unsupported (M8). Tests: expand test_pty_transport.py (validation, result fidelity, dedup, bracketed-paste fidelity, control errors) and test_pty_integration.py (prompt round-trip with newline/leading-slash, fail-loud option). 997 passed, 3 skipped; ruff + mypy clean; one-shot and multi-turn verified live. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Correct the can_use_tool rejection rationale to the verified reason I had described can_use_tool as failing because it "renders a TUI dialog the SDK can't answer / hangs" -- that was unverified (this environment auto-approves all tools, so a permission prompt can't be reproduced here). The actual, code-verifiable reason: can_use_tool is invoked only when transport.read_messages() yields a control_request with subtype can_use_tool (query.py). The PTY transport never produces one, so the callback can never fire regardless of whether the CLI prompts or auto-approves. Tools still run in interactive mode; they just can't be gated through this callback. Update the validation error and docstring accordingly. No behavior change. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add structured question detection/extraction for the PTY transport The interactive TUI renders blocking questions (tool-permission prompts, plan approval, AskUserQuestion forms, app-level confirmations) as on-screen dialogs with no machine-readable signal. Add a detector that reconstructs the rendered screen and parses the dialog into a structured DetectedQuestion. - pty_question.py: dependency-free parser over reconstructed screen lines plus a pyte-backed screen reconstruction helper. Detection is gated on dialog chrome so numbered lists in ordinary model output are not false positives. Extracts kind, tool, target/command, question text, options (with action semantics), preview, and AskUserQuestion tab headers. - pty_cli.py: maintain a pyte screen fed by the drain loop and expose detect_question(); degrades to None when pyte is absent. - pyte is an optional dependency (the pty-introspect extra), mirroring the guarded-import pattern used for opentelemetry. - Unit tests over golden frames for every question kind plus negatives. Extraction was validated against the stream-json can_use_tool payload as ground truth (tool name, target, command, options match) for Write, Bash, Edit, ExitPlanMode, and AskUserQuestion. * PTY: compute cost/usage/model_usage/stop_reason on result (C1, C2, H4, H5) Reconstruct the stream-json ResultMessage data fields from the interactive transcript so the PTY transport is a faithful drop-in: - C2: dedup assistant usage by message.id (transcript writes multiple snapshots of the same message; naive summing over-counted cache tokens) and aggregate the distinct messages of the turn. New TurnUsageAccumulator in _usage.py. - C1: populate total_cost_usd (computed client-side from per-message usage x a small model-pricing table), model_usage, stop_reason, and a non-None permission_denials list. total_cost_usd was silently None, crashing f"${msg.total_cost_usd:.4f}". Pricing validated against the documented stream-json ground truth (within ~0.3% of $0.18327675). - H4: map assistant error/refusal to faithful result subtypes (error_max_turns, error_max_budget_usd, error_during_execution). - H5: num_turns now uses the CLI's per-turn messageCount (the stream-json baseline counts API turns), not a cumulative result counter. - Track the CLI's real session id from transcript records (groundwork for M2). Adds dependency-free unit tests (test_pty_usage.py) and result-fidelity tests. Updated two existing tests that asserted the pre-fix (buggy) num_turns semantics to assert the corrected stream-json behavior. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: populate init/server_info and mcp_status; consistent session_id (C3, M1, M2, partial C4) - C3/M1: build a full init payload from options (tools, mcp_servers, model, permissionMode, apiKeySource, slash_commands, output_style) and use it for both the system/init message and the `initialize` control response, so get_server_info() is no longer empty and the init message is no longer a 5-field stub. The transcript carries no init record, so this is reconstructed from options + faithful defaults. - partial C4: get_mcp_status now returns the configured servers with a `pending` status instead of erroring (no live state is observable over a PTY, so we report configured-but-unknown rather than fabricating "connected"). The genuinely unobservable controls (get_context_usage, rewind_files, stop_task, mcp_toggle) still return a clear error. - M2: track the CLI's real session id from transcript records; result and init use it. Per-message session_id already comes from each record's own sessionId. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: inherit cwd, pass user, restore entrypoint, surface user content (L2, L3, E1, M4) - L2: pass cwd=None to Popen when the caller didn't set options.cwd so the child inherits the parent's working directory (the old transport's behavior), instead of pinning it to a Path.cwd() snapshot. Transcript-path computation still uses the concrete resolved cwd. - L3: pass options.user to Popen so the child runs as the configured OS user, matching the old transport. - E1: restore CLAUDE_CODE_ENTRYPOINT to "sdk-py" (was "sdk-py-pty") so telemetry is keyed identically to the stream-json baseline for drop-in consumers. - M4: surface structured (list) user records even without a tool_result block (e.g. image/document content) instead of dropping them; only the plain-text echo of the typed prompt (string content) is suppressed. Updated the two tests that asserted the pre-fix behavior. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: track live permission mode from transcript records (H1, L5) The CLI writes a permission-mode transcript record whenever the live mode changes. Read it as the source of truth for self._permission_mode, so: - set_permission_mode computes shift+tab steps from the real current mode and is corrected automatically if the CLI's applied mode differs (H1 confirmation, L5 drift), without blocking the call on a transcript round-trip. - set_model tracks the requested model id for follow-up calls. permission-mode records are consumed for state only and not surfaced as SDK messages. Adds a tracking test. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: answer tool-permission dialogs via the TUI detector (C5, C6) The interactive CLI renders tool-permission and plan-approval prompts as on-screen dialogs that block the turn. A background watcher polls the emulated screen (pty_question), and on a blocking dialog: - C5: routes the decision through can_use_tool when configured (mapping allow/deny to the matching dialog option), recording denials on the result's permission_denials list. - C6: with no callback, answers "allow" so the turn completes instead of hanging forever on an unanswered prompt. Answers are sent as the option digit + Enter; each dialog is answered at most once (fingerprint dedup). AskUserQuestion / app dialogs need real content input and are left to the consumer. Consequently can_use_tool is no longer rejected at connect, and the SDK-internal permission_prompt_tool_name="stdio" sentinel (set by the client alongside can_use_tool) is accepted and not forwarded to the interactive CLI as a flag it does not understand. A caller-supplied tool name is still rejected. choose_option() (pure, unit-tested) picks the dialog option for a decision. Adds question-parser and transport-level answering tests. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: honor options.max_buffer_size for the message buffer (M8) max_buffer_size bounded a byte pipe in the old transport; here it bounds the count of buffered message dicts (the closest interactive equivalent), instead of a hardcoded 1000. Updated the validation note accordingly. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: parse structured_output from json_schema turns (H6) When options.output_format is a json_schema format, the CLI constrains the final assistant text to the schema, so parse that text into result["structured_output"] (with a fenced-code-block recovery fallback), mirroring the stream-json baseline. None when no schema is requested or the text is not valid JSON. Unit-tested. Also updates /tmp/dropin_todo.md with the full status of every item. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix PTY drop-in parity: num_turns, assistant-snapshot dedup, calibrated cost C1/C2/H5 were producing wrong values vs live stream-json: - H5 (num_turns): the turn_duration record's `messageCount` counts streamed assistant SNAPSHOTS, not API turns (PONG had messageCount=4 -> wrong num_turns=4). Replaced with `num_turns = (tool_result records) + 1`, derived and verified against 7 live stream-json data points. Live A/B now matches exactly: PONG 1==1, 2-write+2-read 5==5. - C2 (dedup): added emit-side dedup of assistant messages by `message.id` (`_seen_assistant_ids`) so streaming snapshots that share an id but have distinct top-level uuids are not delivered to consumers multiple times. Usage was already deduped via TurnUsageAccumulator. - C1 (cost): recalibrated cache multipliers to live ground truth. The nominal read=0.1 / 1h-write=2.0 over-estimated by ~1.4%; calibrated to read=0.10543, write=1.32145 (fit to 4 live single-call ResultMessage points so the read term cancels). Reproduces all 4 points within 0.06%. - L1: tightened _WARMUP_SECONDS 3.0 -> 1.5 (live-verified the first prompt still submits reliably at 1.5s; even 1.0s worked across repeated runs). Tests: updated test_pty_usage (calibrated rates + live-ground-truth assertion) and test_pty_transport (num_turns = tool_results+1; assistant snapshot emitted once). All gates pass (ruff, mypy, pytest: 1045 passed, 5 skipped). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix model_usage key, cost rates, interrupt, server_info (R1,R2,R4,R5,R6,R8,R9) R1: emit the result usage breakdown under the camelCase wire key `modelUsage` (message_parser reads data.get("modelUsage")) with camelCase per-model sub-keys (inputTokens/outputTokens/cacheReadInputTokens/cacheCreationInputTokens/ webSearchRequests/costUSD/contextWindow/maxOutputTokens). Was snake_case -> ResultMessage.model_usage always None. Live A/B: now populated, identical shape. R2: revert cache multipliers to Anthropic's nominal rates (read 0.1, 5m-write 1.25, 1h-write 2.0). Reproduces the CLI per-model opus costUSD to 0.00% (was +5.5% with the back-solved 0.10543/1.32145, which wrongly fit the result-level total that bundles the unobservable haiku helper line). R4: interrupt() now synthesizes a terminating result (error_during_execution, is_error, result=None) after ESC so receive_response() ends instead of hanging. Track per-turn start time and reset the result latch on each new prompt. R5: add _build_server_info() returning the baseline initialize-control-response key set (account/agents/available_output_styles/commands/models/output_style/ pid), separate from the system/init MESSAGE shape, so get_server_info() consumers find commands/output_style/etc. pid + agents populated; catalogs empty-but-keyed. R6: enrich system/init message: backfill resolved model from the first assistant record; add agents/plugins/skills keys. R8: honor --include-hook-events at the CLI level (was silently dropped). Verified empirically the interactive CLI writes no hook records to the transcript, so hook messages remain unsurfaceable (warning updated). R9: bound _seen_uuids (OrderedDict LRU, cap 2048) to stop unbounded growth in long multi-turn sessions. Adds unit tests for each fix; live A/B verified R1/R2/R4/R5 against the old SDK. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix can_use_tool deny path (RR1,RR2,RR4,RR5) RR1: a can_use_tool deny no longer hangs the turn. The interactive CLI writes no turn_duration after a deny (it goes idle), so _emit_result never fired and receive_response() deadlocked. _answer_question now marks the turn deny-terminated and the tail loop synthesizes a terminating result (subtype=success, is_error=False) once the rejected tool_result lands, mirroring the baseline. Double-emit guarded in _emit_result/_emit_deny_result. RR2: permission_denials now matches the baseline shape {tool_name, tool_use_id, tool_input(full)}. tool_use blocks are indexed by id while tailing; the rejected tool_result is correlated back to its tool_use to recover the real tool_use_id and full original input (the TUI only exposed a target string). RR4: a fresh system/init now leads every turn (was emitted once at connect), matching the baseline's per-turn ordering. RR5: the rejected (permission-denied) tool_result is excluded from num_turns (it is not a real model round-trip); genuine tool errors still count. RR3: verified NOT a divergence -- live A/B shows the baseline also bypasses can_use_tool for allowed_tools-preapproved tools (callback fires only on "ask" decisions), which the PTY already matches. All fixes live-verified against /tmp/venv_old. Unit tests added. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: surface all assistant tool_use/thinking blocks (RV1) The interactive transcript writes each content block of one assistant message as a SEPARATE record sharing one message.id (distinct uuids), and same-id tool_use records are interleaved with the user/tool_result records they trigger. The emit-side _seen_assistant_ids dedup suppressed every same-id record after the first, dropping all tool_use (the agent's actions) and thinking blocks from the consumer stream. Live A/B vs the stream-json baseline shows it emits ONE AssistantMessage per block-record (per-block granularity), NOT a merged message. Emit each assistant block-record directly as its own AssistantMessage; keep usage dedup by message.id and per-uuid true-duplicate dedup. Before/after ToolUseBlock counts (3-file Write, baseline=3): PTY 0 -> 3. Full message sequence and block counts now match the baseline exactly; thinking blocks preserved. No regression to num_turns, cost/usage, the deny path, per-turn init, or duplicate-message suppression. Add TestDedup::test_multiblock_same_id_blocks_all_survive. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix empty-prompt hang, idle interrupt, resume context (RW1,RW2,RW3) RW1: an empty/whitespace prompt cannot be submitted over the TUI, so no turn ran and receive_response() hung forever. _handle_user_message now synthesizes a terminating subtype=success result (mirroring the deny/interrupt synthesis) via _emit_empty_prompt_result so the call returns, matching the baseline subtype. Live A/B: query("")/query(" ") -> subtype=success,is_error=False,num_turns=1 == baseline (was hang). RW2: interrupt() while idle buffered a spurious error_during_execution that the NEXT turn read first, corrupting it. _emit_interrupt_result now no-ops when _turn_start_time is None (no active turn). Live A/B: idle interrupt then query -> subtype=success,result='OK' == baseline (was error_during_execution,result=None). Mid-turn interrupt (R4) unaffected. RW3 (folds in M3): resume/continue did not restore context. - _cli_command no longer auto-appends --session-id over --resume/ --continue unless the caller set one (matches baseline; the auto id made the CLI append to the resumed session while the transport tailed a nonexistent path). - _compute_transcript_path / new _resume_transcript_path resolve the actually-resumed file deterministically. - _tail_loop seeds offset from _initial_tail_offset (resume file size at connect) so the prior turn's turn_duration is not replayed as a stale result. Live A/B: turn1 store 4271, resume in new transport, recall -> '4271' == baseline (was stale 'DONE.'). Unit tests added for all three. No regressions (live A/B): RV1 tool_use x3, num_turns (PONG 1, 3-file 4), usage 1897/5, per-turn init ordering, deny path subtype=success/denials shape -- all match baseline. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add always-on pure-relay API monitor to enrich PTY transport output Interpose a tiny loopback HTTP proxy between the interactive CLI and the Anthropic API so the transport can observe the CLI<->API exchange the .jsonl transcript cannot show. The proxy is ALWAYS ON (no config/env toggle): connect() starts it on 127.0.0.1:<ephemeral>, captures the original ANTHROPIC_BASE_URL as upstream, and points the child CLI's ANTHROPIC_BASE_URL at the proxy; close() stops it. PURE RELAY: request/response bytes are forwarded UNCHANGED (only the Host header is rewritten loopback->upstream, mandatory framing — the API rejects a private-IP Host with 403; body and all other headers are byte-identical). SSE chunks are relayed as they arrive (never buffered), so streaming to the CLI is preserved. Response framing (Content-Length / chunked) is honored so a keep-alive connection cannot hang the relay. The tee is non-fatal: forward first, then parse a gzip/chunk-decoded COPY of /v1/messages behind try/except. Use the captured traffic to enrich the synthesized result: - C1/R3: total_cost_usd / model_usage / usage from the ACTUAL per-call usage across ALL /v1/messages calls in the turn, including the auxiliary helper-model call the transcript never records (no longer under-reporting the helper line). - R7: duration_api_ms = sum of the real per-call request->response durations (no longer equal to the wall-clock duration_ms). - C1: api_error_status surfaced when the turn ends on a non-2xx call (a retried-then-recovered transient is not surfaced). - RV2: can_use_tool receives the FULL tool_use.input (and real tool_use_id) recovered from the intercepted response, correlated by tool name, instead of the scraped {target}. The serve loop uses spawn_detached (not a held-open task group) so it is safe to stop from any task. Falls back to transcript-derived usage and the wall-clock duration when the monitor observes no traffic. Unit tests: tests/test_api_monitor.py (relay transparency vs a fake upstream, SSE tee, gzip, full tool_use input, non-/v1/messages not teed) and TestApiMonitorEnrichment / TestRecoverToolInputRV2 in tests/test_pty_transport.py (cost incl. helper call, summed duration_api_ms, terminal-vs-recovered api_error_status, RV2 full input). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: fix RV2 race, quota over-count, keep-alive, exact endpoint (RL8/RL5/N1/N2/N3) Iteration-2 review fixes for the always-on pure-relay API monitor: RL8 / RV2 (deterministic): _decide_permission now bounded-awaits the relay tee (_await_recovered_tool_input, <=2s/20ms poll) before falling back to the scraped {target}. The CLI renders the permission dialog only after the full /v1/messages response the relay also fully received, so the tee callback is guaranteed within a small window. Live A/B: 5/5 deny runs deliver the full {file_path,content} + real toolu_ id (was nondeterministic). RL5 (cost reconciliation): root-caused the model_usage divergence. The helper/title-gen call DOES traverse the relay; the interactive CLI just sends it with the main model (opus) instead of haiku (small/fast-model resolution differs from the headless baseline) -- a CLI-mode artifact, not a capture gap. Per-call cost math is exact (haiku line matched baseline to the cent). Fixed the real over-count: the synthetic quota_check probe (max_tokens:1, single "quota" user message) is now relayed-but-not-teed so it can't inflate cost/duration/error-status. N1 (keep-alive): _relay_connection serves successive requests on one client connection until close is signaled or framing is ambiguous, so a pooling client never sees a mid-pool dropped connection. N2 (exact endpoint): tee only when urlsplit(path).path == /v1/messages, so /v1/messages/count_tokens and /batches are relayed-but-not-teed while ?beta=true still matches. N3 (body overflow): _relay_request forwards exactly Content-Length body bytes and carries the remainder as leftover for the next pipelined request. Adds unit tests for RV2 await/correlate, exact endpoint match, quota-probe exclusion, and keep-alive/pipelining. Full suite: 1109 passed, 5 skipped. ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: traffic-derived enrichments RL9/RL10/RL11 RL9 (partial messages): when options.include_partial_messages is set, the monitor tee reconstructs stream_event/StreamEvent messages from the SSE events the relay already sees (message_start / content_block_* / message_delta / message_stop), matching the stream-json baseline's StreamEvent shape. Emitted via send_nowait from the monitor serve task (same event loop), best-effort and non-blocking. Nothing is emitted when the option is unset (no behavior change). RL10 (tools catalog): capture the CLI's real resolved tools list + model from the teed /v1/messages request body; init/system message and get_server_info() prefer this observed catalog over the thin options-derived defaults, so they report what the model was actually offered. RL11 (get_context_usage): the previously-unsupported control is now answered from the latest successful per-call usage (input + cache_read + cache_creation tokens => totalTokens), with percentage derived against the model context window. Unobservable breakdowns are returned empty rather than fabricated. Live A/B vs the stream-json baseline confirms each: same StreamEvent kinds, real 29-tool catalog + model on get_server_info, real context token counts. No regression: relay transparency intact, RV2 5/5 full deny input, cost / num_turns / duration_api_ms unchanged. Unit tests added for all three. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: refresh get_server_info catalog, drop stale partial warning (RL10/RL9-warn/N4/N5) RL10: client.get_server_info() now re-issues the `initialize` control request (re-caching _initialization_result) and returns the fresh response instead of the frozen connect-time snapshot, so the live tool catalog/model observed from /v1/messages traffic are surfaced. Safe-if-not-connected; falls back to the cached snapshot on refresh failure. Live: tools 0 (pre-turn) -> 29, model claude-opus-4-8. RL9-warn: deleted the now-false include_partial_messages "has no effect" warning in _validate_options (RL9 emits StreamEvents from the relay-teed SSE stream). The adjacent include_hook_events/stderr warnings remain. N4: documented the PTY/relay limitation in get_server_info()'s docstring (catalog and model become available only after the first turn's traffic; connect-time init is options-derived). N5: kept the non-blocking send_nowait drop in _emit_stream_events (option a) and documented why blocking would stall the relay serve task (transparency violation) -- a necessary residual; buffer is generous and user-tunable. Tests: get_server_info refresh/fallback/not-connected; include_partial_messages no longer warns. Gates: ruff + mypy clean, pytest 1125 passed / 5 skipped. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: map updated_permissions to TUI session-allow (RL12) When can_use_tool returns PermissionResultAllow with non-empty, session-broad updated_permissions, the question-watcher presses the dialog's allow_persist option ("Yes, allow all edits during this session") so the persist intent survives the session, instead of the plain allow_once option. Over-grant guard (_should_persist_allow): press allow_persist only when the dialog actually offers it AND the update is genuinely session-broad (setMode, or an allow addRules/replaceRules scoped to session with no narrowing rule_content). Narrow rules, disk destinations, deny/ask behaviors, or empty updates fall back to allow_once -- never silently grant broader-than-requested via the coarse TUI affordance. Empty updated_permissions -> allow_once, unchanged. Tests: TestPersistAllowRL12 (policy + _decide_permission + _answer_question keystroke + dedup) and choose_option allow_persist cases. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Tighten setMode over-grant guard in _should_persist_allow (RL13) The setMode branch of _should_persist_allow returned True for ANY setMode PermissionUpdate, so allowing a tool while requesting a NARROWING mode (plan/default) would still press the TUI's "allow all edits this session" option -- granting strictly broader than requested, contradicting the guard's "never broader than requested" invariant. Gate the setMode branch on a broadening target: only acceptEdits / bypassPermissions map onto the session accept-edits press; plan / default / None / other fall through to allow_once. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Harden _should_persist_allow into strict positive allowlist (RL14) Close the empty/None-rules over-grant in the addRules/replaceRules branch (rules=[] or None now `continue` instead of vacuously passing the narrowing-rule check and pressing "allow all edits this session"), and restructure the whole function into a strict positive allowlist: it returns True ONLY for (a) setMode to a broadening mode (acceptEdits/bypassPermissions), or (b) addRules/replaceRules with an affirmative behavior=="allow" (None behavior now also falls through), destination in (session, None), non-empty rules, and no narrowing rule_content. Everything else continues to False. Tests: empty-rules add/replaceRules (rules=[]/None x dest session/None) -> allow_once; None behavior -> no persist; unknown update types -> no persist. Existing RL12/RL13 cases re-confirmed. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Require unanimous session-broad list in _should_persist_allow (RL15) Factor the single-element positive test into a static helper _is_session_broad(upd) and stop short-circuiting in _should_persist_allow: it now presses the TUI session-allow option only when EVERY element of updated_permissions is itself session-broad (all(...)), so a mixed list like [setMode acceptEdits, addRules deny] or [broad allow, narrow rule_content] falls back to allow_once instead of dropping the narrowing element. Unanimity is order-independent. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY transport: deliver child stderr via a dedicated pipe (H3) The interactive child was spawned with stdin=stdout=stderr on the slave PTY, so its stderr was muxed onto the terminal and the options.stderr callback never fired (the old stream-json transport delivered stderr per line). Give the child a SEPARATE stderr pipe when options.stderr is set: keep stdin/stdout on the slave PTY (the TUI needs a tty) but route fd 2 to a pipe a background reader splits into lines and hands to the callback, mirroring the old transport's per-line behavior. The reader is non-fatal, and the fd + task are cleaned up in close() (no leak, no orphan). With no callback, stderr stays on the PTY (unchanged). Removed the now-false "stderr not invoked" warning from _validate_options. Live-verified: a real claude turn completes (stdout still on the PTY) and the callback receives the child's real stderr line. Added an end-to-end integration test (fake CLI writes to fd 2; turn still completes and lines are delivered). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY transport: isolate config writes + auto-accept bypass dialog (L4/anthropics#10) L4 -- stop mutating the user's ~/.claude.json. The CLI persists OAuth/login state and settings there and writes onboarding/trust flags into it, so a drop-in must not churn or corrupt it. When the caller/env has NOT set CLAUDE_CONFIG_DIR, default the child to a PERSISTENT SDK-owned dir ($XDG_CACHE_HOME or ~/.cache, under claude-agent-sdk/config) so login survives across runs. Seed it once from the user's real ~/.claude.json (best-effort copy) so auth + settings carry over without ever writing the user's file; _ensure_onboarding_complete then writes the flags into the SDK copy. The same effective config dir feeds transcript-path resolution, so tailing still finds the session .jsonl. An explicit caller/env CLAUDE_CONFIG_DIR is respected unchanged. options.env is never mutated in place. anthropics#10 -- bypassPermissions startup dialog. The screen detector now flags the "Bypass Permissions mode ... 1. No, exit / 2. Yes, I accept" dialog (is_bypass), and the watcher auto-accepts it so the session does not hang. (As root the dialog is suppressed by IS_SANDBOX=1; this covers the non-root case.) Other app dialogs (e.g. folder-trust, handled by config pre-seeding) are left alone. Live-verified: a turn completes against the isolated config dir (auth survived the seeding -- real cost reported) and the user's ~/.claude.json is byte- and mtime-identical before/after; a bypassPermissions turn completes without hanging. Added unit tests for config-dir defaulting/seeding/user-file-untouched/respect, bypass classification, and bypass auto-accept. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add settings-hook IPC bridge + shim for PTY transport win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that together form a deterministic hook channel over the interactive CLI's own settings.json command-hook mechanism. - _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook event. Reads the hook-event JSON on stdin, forwards it (with an auth token) to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no opinion") and exits 0, so a broken shim degrades to the CLI's normal behavior, never a hang/crash. - _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback fallback) that dispatches each event to the user's options.hooks callbacks (matching the old control-protocol hook input/output shapes, incl. async_/ continue_ conversion) and, for PreToolUse when can_use_tool is set, calls can_use_tool and translates PermissionResultAllow(updated_input=...)/ PermissionResultDeny(message=...) into the CLI's hookSpecificOutput.permissionDecision + updatedInput short-circuit. build_hooks_settings() synthesizes the settings.json hooks block wiring the relevant events to the shim (preserving user matchers). Token-authenticated; bounded reads; all dispatch failures return `{}`. Verified end-to-end (server + real shim subprocess): PreToolUse allow with updated_input rewrites the input, deny carries the reason, PostToolUse returns additionalContext, bad-token returns `{}`, user hooks fire with correct input shapes + tool_use_id. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Wire hooks + can_use_tool through the settings-hook IPC bridge (PTY) win #3 (part 2/3): the transport now starts the HookIpcServer in connect(), synthesizes a --settings hooks block wiring the relevant events to the shim, injects the shim's IPC endpoint into the child env, and tears the server down in close(). Programmatic hooks AND can_use_tool now run over this deterministic channel instead of (only) the TUI screen-scraper. pty_cli.py: - connect(): _start_hook_ipc() before _build_command(); env injection of CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the synthesized hooks into an existing --settings JSON object (or appends one; non-object existing value -> second --settings, which the CLI merges). - _validate_options(): removed the `hooks` rejection (now supported). - _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the transport falls back to the existing watcher and hooks simply don't fire. - close(): stop the hook IPC server (cancel serve task + unlink socket; no leak). - Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the watcher (the PreToolUse hook already decided server-side and the CLI short-circuits the dialog). _on_hook_permission_decision records a hook-channel deny into _pending_denied_tools so result.permission_denials still correlates. _hook_ipc.py: added the optional on_permission_decision sink (fires after a can_use_tool PreToolUse decision) so the transport can record hook-channel denies. Tests: - test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process dispatch (user hook shape, matcher filtering, async_/continue_ conversion, can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip, bad-token and no-endpoint fail-open. - test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI that actually runs the shim): PreToolUse fires with correct input shape + tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool allow with updated_input ACTUALLY rewrites the executed input (content -> REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response. - test_pty_integration validation test flipped: hooks accepted (not rejected); permission_prompt_tool_name still rejected. - test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge) + TestCanUseToolViaHookGuard (no double-invoke; deny recorded). Full suite 1220 passed / 5 skipped. ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Document hooks + can_use_tool support; live A/B verification win #3 (part 3/3): update the PTY transport module docstring to reflect that programmatic hooks and can_use_tool are now supported via the settings-hook IPC bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher is the fallback for when the bridge can't start (and always handles plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks). Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`: - updated_input ACTUALLY applied: a Write the model issued with content="ORIGINAL_MODEL_CONTENT" was rewritten by can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"}); the file on disk contained the REWRITTEN content (PostToolUse tool_response confirmed it). Baseline old SDK identical. - can_use_tool deny via the hook channel: tool not run, full input + real toolu_ id delivered, result.permission_denials carries the baseline shape. - PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort, hook_event_name, permission_mode, session_id, tool_input, tool_name, tool_use_id, transcript_path]; PostToolUse fires with tool_response. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix hook-IPC review findings W1-W4 (PTY transport) W1: gate the Unix-socket branch on hasattr(socket, "AF_UNIX") instead of the always-False hasattr(os, "AF_UNIX"), so POSIX uses a 0700-dir Unix domain socket (TCP loopback stays the fallback). W2: close the can_use_tool permission bypass. build_hooks_settings now ALWAYS emits a catch-all PreToolUse shim entry when can_use_tool is set (in addition to any narrow user matchers), so the callback is consulted for every tool. The IPC dispatch dedups can_use_tool per tool_use_id so a tool matching both the catch-all and the user's narrow matcher invokes the callback only once. W3: _close_stderr_pipe now closes both the write and read pipe ends (and clears _stderr_read_fd) on every failed-spawn path, so a direct caller that skips close() after a failed connect() does not leak the read fd. W4: _matching_callbacks fires non-tool-event hooks (UserPromptSubmit/Stop/...) regardless of the matcher pattern, mirroring the CLI's getMatchingHooks (undefined matchQuery -> all matchers). Settings synthesis already omits the tool matcher for non-tool events. Tests: +catch-all/dedup/unmatched-tool/no-double-invoke (W2), unix-socket+0700 (W1), failed-spawn fd cleanup (W3), non-tool-event matcher fires (W4). Gates green: ruff + mypy clean, pytest 1234 passed / 5 skipped. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix hook-IPC permission merge: deny-wins + per-id concurrency (W5/W6) W5 (MEDIUM): a can_use_tool ALLOW could silently overwrite a user PreToolUse hook DENY and leak a stale "hook-deny" reason into the allow (self-contradictory, under-enforced). _dispatch now routes the decision through HookIpcServer._merge_permission_decision: deny from EITHER source wins, and the permission triplet (permissionDecision / permissionDecisionReason / updatedInput) is replaced as an atomic unit so no stale field leaks. Non-permission hookSpecificOutput fields are kept. updated_input still applies on a both-allow result. W6 (LOW-MED): _run_can_use_tool_deduped held _cut_lock across the user can_use_tool callback, serializing distinct concurrent tool_use_ids (and risking deadlock on re-entrancy). Switched to double-checked locking with a per-tool_use_id _CutInflight marker: the lock now only guards the cache /in-flight dicts, distinct ids run the callback concurrently, and a duplicate fire for the same id awaits+replays the single result (exactly-once consult). perm is pre-bound so a raising callback cannot hang waiters. Bounded 512-entry FIFO cache retained. Tests: W5 deny-wins matrix (reason correctness + updated_input-on-allow + non-perm-field preservation); W6 distinct-id non-serialization, same-id single-consult, and raising-callback-wakes-waiters. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH --------- Co-authored-by: Claude <noreply@anthropic.com>
bddppq
referenced
this pull request
in bddppq/claude-agent-sdk-python
Jun 9, 2026
…ment) (#1) * Run the CLI in interactive mode over a PTY instead of headless pipes Replace the default transport with a PTY-based one that drives the *interactive* Claude Code CLI rather than the headless stream-json pipe (formerly `-p`/`--print`). The CLI is launched attached to a pseudo-terminal, prompts are typed in as real keystrokes, and responses are read by tailing the session `.jsonl` transcript and translating its records back into the SDK message dicts the parser already understands. This is invisible to callers: the public `query()` / `ClaudeSDKClient` API is unchanged. Details: - New `PtyCLITransport` (POSIX only): allocates a PTY in raw mode, spawns the interactive CLI in its own process group, drains terminal output, and tails `<config>/projects/<sanitized-cwd>/<session-id>.jsonl`. - Reuses `SubprocessCLITransport`'s command/option logic, stripping the stream-json I/O flags that only apply to headless mode. - Pre-seeds `hasCompletedOnboarding` and per-folder `hasTrustDialogAccepted` in the CLI config so the TUI drops straight to the prompt, and normalizes `IS_SANDBOX=1` when running as root. - Acknowledges the initialize control request locally and synthesizes a `result` from the transcript's `turn_duration` record so existing one-shot and multi-turn flows terminate correctly. Interrupt types ESC. - Surfaces an error result if the CLI exits early instead of ending silently. - `SubprocessCLITransport` is retained for explicit/custom use. - Adds unit tests for translation/command/env/keystroke logic, repoints existing transport mocks to the new default, and adds an example. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Make the interactive PTY transport the SDK's only transport Fully replace the headless subprocess `-p`/stream-json transport with the interactive PTY transport, and map every SDK control operation to its interactive-TUI equivalent so the rest of the SDK keeps working. Transport replacement: - Extract the transport-agnostic logic (CLI discovery, option->flag command building, subprocess env + OTEL propagation, version check) into a shared `_cli_command` module. The command it builds is the interactive command: no stream-json I/O flags, no `--print`, and always `--session-id` so the transcript can be located. - `PtyCLITransport` is now self-contained and the sole transport; delete `subprocess_cli.py`. Control operations -> interactive equivalents (PtyCLITransport): - interrupt -> ESC - set_permission_mode -> shift+tab cycles default/acceptEdits/plan - set_model -> `/model` slash command - mcp_status / get_context_usage / rewind_files -> `/mcp` `/context` `/rewind` - initialize and other control requests are acknowledged locally so the handshake proceeds. Tests: - Add `test_cli_command.py` covering command building, env (+OTEL, IS_SANDBOX), skills/sandbox/settings, discovery, and version check. - Migrate transport mocks/patches to the new default; drop tests that were specific to the removed stream-json pipe (stdin streaming, stdout buffering, `--session-mirror` flag). - Whole suite passes (943 passed, 3 skipped); ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add comprehensive drop-in tests for the interactive PTY transport Prove the new transport behaves as a seamless replacement for the former subprocess `-p` transport, exercised through the unchanged public API. - test_pty_integration.py: end-to-end tests that spawn the real PtyCLITransport (PTY allocation, keystroke injection, transcript tailing, message translation) against a fake `claude` binary and drive it through `query()` and `ClaudeSDKClient`. Covers one-shot, multi-turn, tool-result surfacing, plain-prompt suppression, and early-exit error reporting. Runs under both asyncio and trio; no auth or real model needed. - test_pty_transport.py: add deterministic unit tests for the transcript tail loop (translation, bookkeeping-record skipping, malformed lines, tool results), early-exit error result, and the control-request -> interactive mappings (permission-mode shift+tab cycling incl. wrap-around and the non-cyclable bypass case, set_model `/model`, `/mcp` `/context` `/rewind`), plus init/control-response ordering. - pty_cli: extract the TUI warmup delay into a module-level `_WARMUP_SECONDS` so tests can shrink it; no behavior change in production. Full suite: 972 passed, 3 skipped; ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Harden the interactive PTY transport against the review findings Address the breakage cases surfaced by the nitpick review: fix what is fixable, and convert the architecturally-unfixable cases (features that need the bidirectional control channel the interactive CLI does not expose) from silent no-ops / hangs into loud, actionable errors. Fail loud instead of silently breaking (C1/C2/C3/C5, H9): - _validate_options() rejects can_use_tool, hooks, in-process SDK MCP servers (external MCP still works), session_store, and permission_prompt_tool_name with a clear CLIConnectionError at connect(); include_partial_messages / include_hook_events / stderr warn rather than silently doing nothing. - Unsupported control requests (mcp_status, get_context_usage, mcp_reconnect, mcp_toggle, stop_task, rewind_files, uncyclable permission modes) now return an error control_response instead of a fake success. Correctness fixes: - Prompt fidelity (H4/H5/H6): type prompts via bracketed paste so newlines and special characters are preserved; prepend a single space for a leading "/"/"!"/"#" so the TUI doesn't enter command mode; warn on dropped non-text (multimodal) blocks. - Result fidelity (H1/H2/H3): synthesize the result statefully from the turn -- final assistant text, summed usage, real num_turns, and is_error on refusal/error (total_cost_usd remains unavailable from the transcript). - Concurrency (M1): serialize all PTY keystroke sequences with a write lock. - Transcript discovery (L2/L3): resolve the path lazily with safe fallbacks (exact path -> our-session-id rglob -> newest in our project dir, gated on fork/resume and scoped so a concurrent unrelated session is never picked up). - Compaction tolerance (L4): re-read on truncation, dedup records by uuid. - Parser robustness (L1): coerce string/missing content and backfill model. - Lifecycle (L6): atexit cleanup of child process groups + master fds; atomic ~/.claude.json update (temp file + os.replace) to avoid clobbering (H10). - Carry parent_tool_use_id from the transcript (M9). Documented limitations: POSIX-only (M7), stderr/max_buffer_size inert (M4/M5), structured_output/total_cost_usd unavailable (M6), rewind-by-id unsupported (M8). Tests: expand test_pty_transport.py (validation, result fidelity, dedup, bracketed-paste fidelity, control errors) and test_pty_integration.py (prompt round-trip with newline/leading-slash, fail-loud option). 997 passed, 3 skipped; ruff + mypy clean; one-shot and multi-turn verified live. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Correct the can_use_tool rejection rationale to the verified reason I had described can_use_tool as failing because it "renders a TUI dialog the SDK can't answer / hangs" -- that was unverified (this environment auto-approves all tools, so a permission prompt can't be reproduced here). The actual, code-verifiable reason: can_use_tool is invoked only when transport.read_messages() yields a control_request with subtype can_use_tool (query.py). The PTY transport never produces one, so the callback can never fire regardless of whether the CLI prompts or auto-approves. Tools still run in interactive mode; they just can't be gated through this callback. Update the validation error and docstring accordingly. No behavior change. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add structured question detection/extraction for the PTY transport The interactive TUI renders blocking questions (tool-permission prompts, plan approval, AskUserQuestion forms, app-level confirmations) as on-screen dialogs with no machine-readable signal. Add a detector that reconstructs the rendered screen and parses the dialog into a structured DetectedQuestion. - pty_question.py: dependency-free parser over reconstructed screen lines plus a pyte-backed screen reconstruction helper. Detection is gated on dialog chrome so numbered lists in ordinary model output are not false positives. Extracts kind, tool, target/command, question text, options (with action semantics), preview, and AskUserQuestion tab headers. - pty_cli.py: maintain a pyte screen fed by the drain loop and expose detect_question(); degrades to None when pyte is absent. - pyte is an optional dependency (the pty-introspect extra), mirroring the guarded-import pattern used for opentelemetry. - Unit tests over golden frames for every question kind plus negatives. Extraction was validated against the stream-json can_use_tool payload as ground truth (tool name, target, command, options match) for Write, Bash, Edit, ExitPlanMode, and AskUserQuestion. * PTY: compute cost/usage/model_usage/stop_reason on result (C1, C2, H4, H5) Reconstruct the stream-json ResultMessage data fields from the interactive transcript so the PTY transport is a faithful drop-in: - C2: dedup assistant usage by message.id (transcript writes multiple snapshots of the same message; naive summing over-counted cache tokens) and aggregate the distinct messages of the turn. New TurnUsageAccumulator in _usage.py. - C1: populate total_cost_usd (computed client-side from per-message usage x a small model-pricing table), model_usage, stop_reason, and a non-None permission_denials list. total_cost_usd was silently None, crashing f"${msg.total_cost_usd:.4f}". Pricing validated against the documented stream-json ground truth (within ~0.3% of $0.18327675). - H4: map assistant error/refusal to faithful result subtypes (error_max_turns, error_max_budget_usd, error_during_execution). - H5: num_turns now uses the CLI's per-turn messageCount (the stream-json baseline counts API turns), not a cumulative result counter. - Track the CLI's real session id from transcript records (groundwork for M2). Adds dependency-free unit tests (test_pty_usage.py) and result-fidelity tests. Updated two existing tests that asserted the pre-fix (buggy) num_turns semantics to assert the corrected stream-json behavior. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: populate init/server_info and mcp_status; consistent session_id (C3, M1, M2, partial C4) - C3/M1: build a full init payload from options (tools, mcp_servers, model, permissionMode, apiKeySource, slash_commands, output_style) and use it for both the system/init message and the `initialize` control response, so get_server_info() is no longer empty and the init message is no longer a 5-field stub. The transcript carries no init record, so this is reconstructed from options + faithful defaults. - partial C4: get_mcp_status now returns the configured servers with a `pending` status instead of erroring (no live state is observable over a PTY, so we report configured-but-unknown rather than fabricating "connected"). The genuinely unobservable controls (get_context_usage, rewind_files, stop_task, mcp_toggle) still return a clear error. - M2: track the CLI's real session id from transcript records; result and init use it. Per-message session_id already comes from each record's own sessionId. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: inherit cwd, pass user, restore entrypoint, surface user content (L2, L3, E1, M4) - L2: pass cwd=None to Popen when the caller didn't set options.cwd so the child inherits the parent's working directory (the old transport's behavior), instead of pinning it to a Path.cwd() snapshot. Transcript-path computation still uses the concrete resolved cwd. - L3: pass options.user to Popen so the child runs as the configured OS user, matching the old transport. - E1: restore CLAUDE_CODE_ENTRYPOINT to "sdk-py" (was "sdk-py-pty") so telemetry is keyed identically to the stream-json baseline for drop-in consumers. - M4: surface structured (list) user records even without a tool_result block (e.g. image/document content) instead of dropping them; only the plain-text echo of the typed prompt (string content) is suppressed. Updated the two tests that asserted the pre-fix behavior. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: track live permission mode from transcript records (H1, L5) The CLI writes a permission-mode transcript record whenever the live mode changes. Read it as the source of truth for self._permission_mode, so: - set_permission_mode computes shift+tab steps from the real current mode and is corrected automatically if the CLI's applied mode differs (H1 confirmation, L5 drift), without blocking the call on a transcript round-trip. - set_model tracks the requested model id for follow-up calls. permission-mode records are consumed for state only and not surfaced as SDK messages. Adds a tracking test. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: answer tool-permission dialogs via the TUI detector (C5, C6) The interactive CLI renders tool-permission and plan-approval prompts as on-screen dialogs that block the turn. A background watcher polls the emulated screen (pty_question), and on a blocking dialog: - C5: routes the decision through can_use_tool when configured (mapping allow/deny to the matching dialog option), recording denials on the result's permission_denials list. - C6: with no callback, answers "allow" so the turn completes instead of hanging forever on an unanswered prompt. Answers are sent as the option digit + Enter; each dialog is answered at most once (fingerprint dedup). AskUserQuestion / app dialogs need real content input and are left to the consumer. Consequently can_use_tool is no longer rejected at connect, and the SDK-internal permission_prompt_tool_name="stdio" sentinel (set by the client alongside can_use_tool) is accepted and not forwarded to the interactive CLI as a flag it does not understand. A caller-supplied tool name is still rejected. choose_option() (pure, unit-tested) picks the dialog option for a decision. Adds question-parser and transport-level answering tests. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: honor options.max_buffer_size for the message buffer (M8) max_buffer_size bounded a byte pipe in the old transport; here it bounds the count of buffered message dicts (the closest interactive equivalent), instead of a hardcoded 1000. Updated the validation note accordingly. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY: parse structured_output from json_schema turns (H6) When options.output_format is a json_schema format, the CLI constrains the final assistant text to the schema, so parse that text into result["structured_output"] (with a fenced-code-block recovery fallback), mirroring the stream-json baseline. None when no schema is requested or the text is not valid JSON. Unit-tested. Also updates /tmp/dropin_todo.md with the full status of every item. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix PTY drop-in parity: num_turns, assistant-snapshot dedup, calibrated cost C1/C2/H5 were producing wrong values vs live stream-json: - H5 (num_turns): the turn_duration record's `messageCount` counts streamed assistant SNAPSHOTS, not API turns (PONG had messageCount=4 -> wrong num_turns=4). Replaced with `num_turns = (tool_result records) + 1`, derived and verified against 7 live stream-json data points. Live A/B now matches exactly: PONG 1==1, 2-write+2-read 5==5. - C2 (dedup): added emit-side dedup of assistant messages by `message.id` (`_seen_assistant_ids`) so streaming snapshots that share an id but have distinct top-level uuids are not delivered to consumers multiple times. Usage was already deduped via TurnUsageAccumulator. - C1 (cost): recalibrated cache multipliers to live ground truth. The nominal read=0.1 / 1h-write=2.0 over-estimated by ~1.4%; calibrated to read=0.10543, write=1.32145 (fit to 4 live single-call ResultMessage points so the read term cancels). Reproduces all 4 points within 0.06%. - L1: tightened _WARMUP_SECONDS 3.0 -> 1.5 (live-verified the first prompt still submits reliably at 1.5s; even 1.0s worked across repeated runs). Tests: updated test_pty_usage (calibrated rates + live-ground-truth assertion) and test_pty_transport (num_turns = tool_results+1; assistant snapshot emitted once). All gates pass (ruff, mypy, pytest: 1045 passed, 5 skipped). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix model_usage key, cost rates, interrupt, server_info (R1,R2,R4,R5,R6,R8,R9) R1: emit the result usage breakdown under the camelCase wire key `modelUsage` (message_parser reads data.get("modelUsage")) with camelCase per-model sub-keys (inputTokens/outputTokens/cacheReadInputTokens/cacheCreationInputTokens/ webSearchRequests/costUSD/contextWindow/maxOutputTokens). Was snake_case -> ResultMessage.model_usage always None. Live A/B: now populated, identical shape. R2: revert cache multipliers to Anthropic's nominal rates (read 0.1, 5m-write 1.25, 1h-write 2.0). Reproduces the CLI per-model opus costUSD to 0.00% (was +5.5% with the back-solved 0.10543/1.32145, which wrongly fit the result-level total that bundles the unobservable haiku helper line). R4: interrupt() now synthesizes a terminating result (error_during_execution, is_error, result=None) after ESC so receive_response() ends instead of hanging. Track per-turn start time and reset the result latch on each new prompt. R5: add _build_server_info() returning the baseline initialize-control-response key set (account/agents/available_output_styles/commands/models/output_style/ pid), separate from the system/init MESSAGE shape, so get_server_info() consumers find commands/output_style/etc. pid + agents populated; catalogs empty-but-keyed. R6: enrich system/init message: backfill resolved model from the first assistant record; add agents/plugins/skills keys. R8: honor --include-hook-events at the CLI level (was silently dropped). Verified empirically the interactive CLI writes no hook records to the transcript, so hook messages remain unsurfaceable (warning updated). R9: bound _seen_uuids (OrderedDict LRU, cap 2048) to stop unbounded growth in long multi-turn sessions. Adds unit tests for each fix; live A/B verified R1/R2/R4/R5 against the old SDK. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix can_use_tool deny path (RR1,RR2,RR4,RR5) RR1: a can_use_tool deny no longer hangs the turn. The interactive CLI writes no turn_duration after a deny (it goes idle), so _emit_result never fired and receive_response() deadlocked. _answer_question now marks the turn deny-terminated and the tail loop synthesizes a terminating result (subtype=success, is_error=False) once the rejected tool_result lands, mirroring the baseline. Double-emit guarded in _emit_result/_emit_deny_result. RR2: permission_denials now matches the baseline shape {tool_name, tool_use_id, tool_input(full)}. tool_use blocks are indexed by id while tailing; the rejected tool_result is correlated back to its tool_use to recover the real tool_use_id and full original input (the TUI only exposed a target string). RR4: a fresh system/init now leads every turn (was emitted once at connect), matching the baseline's per-turn ordering. RR5: the rejected (permission-denied) tool_result is excluded from num_turns (it is not a real model round-trip); genuine tool errors still count. RR3: verified NOT a divergence -- live A/B shows the baseline also bypasses can_use_tool for allowed_tools-preapproved tools (callback fires only on "ask" decisions), which the PTY already matches. All fixes live-verified against /tmp/venv_old. Unit tests added. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: surface all assistant tool_use/thinking blocks (RV1) The interactive transcript writes each content block of one assistant message as a SEPARATE record sharing one message.id (distinct uuids), and same-id tool_use records are interleaved with the user/tool_result records they trigger. The emit-side _seen_assistant_ids dedup suppressed every same-id record after the first, dropping all tool_use (the agent's actions) and thinking blocks from the consumer stream. Live A/B vs the stream-json baseline shows it emits ONE AssistantMessage per block-record (per-block granularity), NOT a merged message. Emit each assistant block-record directly as its own AssistantMessage; keep usage dedup by message.id and per-uuid true-duplicate dedup. Before/after ToolUseBlock counts (3-file Write, baseline=3): PTY 0 -> 3. Full message sequence and block counts now match the baseline exactly; thinking blocks preserved. No regression to num_turns, cost/usage, the deny path, per-turn init, or duplicate-message suppression. Add TestDedup::test_multiblock_same_id_blocks_all_survive. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY drop-in: fix empty-prompt hang, idle interrupt, resume context (RW1,RW2,RW3) RW1: an empty/whitespace prompt cannot be submitted over the TUI, so no turn ran and receive_response() hung forever. _handle_user_message now synthesizes a terminating subtype=success result (mirroring the deny/interrupt synthesis) via _emit_empty_prompt_result so the call returns, matching the baseline subtype. Live A/B: query("")/query(" ") -> subtype=success,is_error=False,num_turns=1 == baseline (was hang). RW2: interrupt() while idle buffered a spurious error_during_execution that the NEXT turn read first, corrupting it. _emit_interrupt_result now no-ops when _turn_start_time is None (no active turn). Live A/B: idle interrupt then query -> subtype=success,result='OK' == baseline (was error_during_execution,result=None). Mid-turn interrupt (R4) unaffected. RW3 (folds in M3): resume/continue did not restore context. - _cli_command no longer auto-appends --session-id over --resume/ --continue unless the caller set one (matches baseline; the auto id made the CLI append to the resumed session while the transport tailed a nonexistent path). - _compute_transcript_path / new _resume_transcript_path resolve the actually-resumed file deterministically. - _tail_loop seeds offset from _initial_tail_offset (resume file size at connect) so the prior turn's turn_duration is not replayed as a stale result. Live A/B: turn1 store 4271, resume in new transport, recall -> '4271' == baseline (was stale 'DONE.'). Unit tests added for all three. No regressions (live A/B): RV1 tool_use x3, num_turns (PONG 1, 3-file 4), usage 1897/5, per-turn init ordering, deny path subtype=success/denials shape -- all match baseline. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add always-on pure-relay API monitor to enrich PTY transport output Interpose a tiny loopback HTTP proxy between the interactive CLI and the Anthropic API so the transport can observe the CLI<->API exchange the .jsonl transcript cannot show. The proxy is ALWAYS ON (no config/env toggle): connect() starts it on 127.0.0.1:<ephemeral>, captures the original ANTHROPIC_BASE_URL as upstream, and points the child CLI's ANTHROPIC_BASE_URL at the proxy; close() stops it. PURE RELAY: request/response bytes are forwarded UNCHANGED (only the Host header is rewritten loopback->upstream, mandatory framing — the API rejects a private-IP Host with 403; body and all other headers are byte-identical). SSE chunks are relayed as they arrive (never buffered), so streaming to the CLI is preserved. Response framing (Content-Length / chunked) is honored so a keep-alive connection cannot hang the relay. The tee is non-fatal: forward first, then parse a gzip/chunk-decoded COPY of /v1/messages behind try/except. Use the captured traffic to enrich the synthesized result: - C1/R3: total_cost_usd / model_usage / usage from the ACTUAL per-call usage across ALL /v1/messages calls in the turn, including the auxiliary helper-model call the transcript never records (no longer under-reporting the helper line). - R7: duration_api_ms = sum of the real per-call request->response durations (no longer equal to the wall-clock duration_ms). - C1: api_error_status surfaced when the turn ends on a non-2xx call (a retried-then-recovered transient is not surfaced). - RV2: can_use_tool receives the FULL tool_use.input (and real tool_use_id) recovered from the intercepted response, correlated by tool name, instead of the scraped {target}. The serve loop uses spawn_detached (not a held-open task group) so it is safe to stop from any task. Falls back to transcript-derived usage and the wall-clock duration when the monitor observes no traffic. Unit tests: tests/test_api_monitor.py (relay transparency vs a fake upstream, SSE tee, gzip, full tool_use input, non-/v1/messages not teed) and TestApiMonitorEnrichment / TestRecoverToolInputRV2 in tests/test_pty_transport.py (cost incl. helper call, summed duration_api_ms, terminal-vs-recovered api_error_status, RV2 full input). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: fix RV2 race, quota over-count, keep-alive, exact endpoint (RL8/RL5/N1/N2/N3) Iteration-2 review fixes for the always-on pure-relay API monitor: RL8 / RV2 (deterministic): _decide_permission now bounded-awaits the relay tee (_await_recovered_tool_input, <=2s/20ms poll) before falling back to the scraped {target}. The CLI renders the permission dialog only after the full /v1/messages response the relay also fully received, so the tee callback is guaranteed within a small window. Live A/B: 5/5 deny runs deliver the full {file_path,content} + real toolu_ id (was nondeterministic). RL5 (cost reconciliation): root-caused the model_usage divergence. The helper/title-gen call DOES traverse the relay; the interactive CLI just sends it with the main model (opus) instead of haiku (small/fast-model resolution differs from the headless baseline) -- a CLI-mode artifact, not a capture gap. Per-call cost math is exact (haiku line matched baseline to the cent). Fixed the real over-count: the synthetic quota_check probe (max_tokens:1, single "quota" user message) is now relayed-but-not-teed so it can't inflate cost/duration/error-status. N1 (keep-alive): _relay_connection serves successive requests on one client connection until close is signaled or framing is ambiguous, so a pooling client never sees a mid-pool dropped connection. N2 (exact endpoint): tee only when urlsplit(path).path == /v1/messages, so /v1/messages/count_tokens and /batches are relayed-but-not-teed while ?beta=true still matches. N3 (body overflow): _relay_request forwards exactly Content-Length body bytes and carries the remainder as leftover for the next pipelined request. Adds unit tests for RV2 await/correlate, exact endpoint match, quota-probe exclusion, and keep-alive/pipelining. Full suite: 1109 passed, 5 skipped. ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: traffic-derived enrichments RL9/RL10/RL11 RL9 (partial messages): when options.include_partial_messages is set, the monitor tee reconstructs stream_event/StreamEvent messages from the SSE events the relay already sees (message_start / content_block_* / message_delta / message_stop), matching the stream-json baseline's StreamEvent shape. Emitted via send_nowait from the monitor serve task (same event loop), best-effort and non-blocking. Nothing is emitted when the option is unset (no behavior change). RL10 (tools catalog): capture the CLI's real resolved tools list + model from the teed /v1/messages request body; init/system message and get_server_info() prefer this observed catalog over the thin options-derived defaults, so they report what the model was actually offered. RL11 (get_context_usage): the previously-unsupported control is now answered from the latest successful per-call usage (input + cache_read + cache_creation tokens => totalTokens), with percentage derived against the model context window. Unobservable breakdowns are returned empty rather than fabricated. Live A/B vs the stream-json baseline confirms each: same StreamEvent kinds, real 29-tool catalog + model on get_server_info, real context token counts. No regression: relay transparency intact, RV2 5/5 full deny input, cost / num_turns / duration_api_ms unchanged. Unit tests added for all three. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: refresh get_server_info catalog, drop stale partial warning (RL10/RL9-warn/N4/N5) RL10: client.get_server_info() now re-issues the `initialize` control request (re-caching _initialization_result) and returns the fresh response instead of the frozen connect-time snapshot, so the live tool catalog/model observed from /v1/messages traffic are surfaced. Safe-if-not-connected; falls back to the cached snapshot on refresh failure. Live: tools 0 (pre-turn) -> 29, model claude-opus-4-8. RL9-warn: deleted the now-false include_partial_messages "has no effect" warning in _validate_options (RL9 emits StreamEvents from the relay-teed SSE stream). The adjacent include_hook_events/stderr warnings remain. N4: documented the PTY/relay limitation in get_server_info()'s docstring (catalog and model become available only after the first turn's traffic; connect-time init is options-derived). N5: kept the non-blocking send_nowait drop in _emit_stream_events (option a) and documented why blocking would stall the relay serve task (transparency violation) -- a necessary residual; buffer is generous and user-tunable. Tests: get_server_info refresh/fallback/not-connected; include_partial_messages no longer warns. Gates: ruff + mypy clean, pytest 1125 passed / 5 skipped. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY relay: map updated_permissions to TUI session-allow (RL12) When can_use_tool returns PermissionResultAllow with non-empty, session-broad updated_permissions, the question-watcher presses the dialog's allow_persist option ("Yes, allow all edits during this session") so the persist intent survives the session, instead of the plain allow_once option. Over-grant guard (_should_persist_allow): press allow_persist only when the dialog actually offers it AND the update is genuinely session-broad (setMode, or an allow addRules/replaceRules scoped to session with no narrowing rule_content). Narrow rules, disk destinations, deny/ask behaviors, or empty updates fall back to allow_once -- never silently grant broader-than-requested via the coarse TUI affordance. Empty updated_permissions -> allow_once, unchanged. Tests: TestPersistAllowRL12 (policy + _decide_permission + _answer_question keystroke + dedup) and choose_option allow_persist cases. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Tighten setMode over-grant guard in _should_persist_allow (RL13) The setMode branch of _should_persist_allow returned True for ANY setMode PermissionUpdate, so allowing a tool while requesting a NARROWING mode (plan/default) would still press the TUI's "allow all edits this session" option -- granting strictly broader than requested, contradicting the guard's "never broader than requested" invariant. Gate the setMode branch on a broadening target: only acceptEdits / bypassPermissions map onto the session accept-edits press; plan / default / None / other fall through to allow_once. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Harden _should_persist_allow into strict positive allowlist (RL14) Close the empty/None-rules over-grant in the addRules/replaceRules branch (rules=[] or None now `continue` instead of vacuously passing the narrowing-rule check and pressing "allow all edits this session"), and restructure the whole function into a strict positive allowlist: it returns True ONLY for (a) setMode to a broadening mode (acceptEdits/bypassPermissions), or (b) addRules/replaceRules with an affirmative behavior=="allow" (None behavior now also falls through), destination in (session, None), non-empty rules, and no narrowing rule_content. Everything else continues to False. Tests: empty-rules add/replaceRules (rules=[]/None x dest session/None) -> allow_once; None behavior -> no persist; unknown update types -> no persist. Existing RL12/RL13 cases re-confirmed. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Require unanimous session-broad list in _should_persist_allow (RL15) Factor the single-element positive test into a static helper _is_session_broad(upd) and stop short-circuiting in _should_persist_allow: it now presses the TUI session-allow option only when EVERY element of updated_permissions is itself session-broad (all(...)), so a mixed list like [setMode acceptEdits, addRules deny] or [broad allow, narrow rule_content] falls back to allow_once instead of dropping the narrowing element. Unanimity is order-independent. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY transport: deliver child stderr via a dedicated pipe (H3) The interactive child was spawned with stdin=stdout=stderr on the slave PTY, so its stderr was muxed onto the terminal and the options.stderr callback never fired (the old stream-json transport delivered stderr per line). Give the child a SEPARATE stderr pipe when options.stderr is set: keep stdin/stdout on the slave PTY (the TUI needs a tty) but route fd 2 to a pipe a background reader splits into lines and hands to the callback, mirroring the old transport's per-line behavior. The reader is non-fatal, and the fd + task are cleaned up in close() (no leak, no orphan). With no callback, stderr stays on the PTY (unchanged). Removed the now-false "stderr not invoked" warning from _validate_options. Live-verified: a real claude turn completes (stdout still on the PTY) and the callback receives the child's real stderr line. Added an end-to-end integration test (fake CLI writes to fd 2; turn still completes and lines are delivered). https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * PTY transport: isolate config writes + auto-accept bypass dialog (L4/anthropics#10) L4 -- stop mutating the user's ~/.claude.json. The CLI persists OAuth/login state and settings there and writes onboarding/trust flags into it, so a drop-in must not churn or corrupt it. When the caller/env has NOT set CLAUDE_CONFIG_DIR, default the child to a PERSISTENT SDK-owned dir ($XDG_CACHE_HOME or ~/.cache, under claude-agent-sdk/config) so login survives across runs. Seed it once from the user's real ~/.claude.json (best-effort copy) so auth + settings carry over without ever writing the user's file; _ensure_onboarding_complete then writes the flags into the SDK copy. The same effective config dir feeds transcript-path resolution, so tailing still finds the session .jsonl. An explicit caller/env CLAUDE_CONFIG_DIR is respected unchanged. options.env is never mutated in place. anthropics#10 -- bypassPermissions startup dialog. The screen detector now flags the "Bypass Permissions mode ... 1. No, exit / 2. Yes, I accept" dialog (is_bypass), and the watcher auto-accepts it so the session does not hang. (As root the dialog is suppressed by IS_SANDBOX=1; this covers the non-root case.) Other app dialogs (e.g. folder-trust, handled by config pre-seeding) are left alone. Live-verified: a turn completes against the isolated config dir (auth survived the seeding -- real cost reported) and the user's ~/.claude.json is byte- and mtime-identical before/after; a bypassPermissions turn completes without hanging. Added unit tests for config-dir defaulting/seeding/user-file-untouched/respect, bypass classification, and bypass auto-accept. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Add settings-hook IPC bridge + shim for PTY transport win #3 (part 1/3): the SDK-process IPC server and the CLI-side shim that together form a deterministic hook channel over the interactive CLI's own settings.json command-hook mechanism. - _hook_shim.py: tiny, stdlib-only `python -m` command the CLI runs per hook event. Reads the hook-event JSON on stdin, forwards it (with an auth token) to the SDK's IPC endpoint via CLAUDE_AGENT_SDK_HOOK_IPC, writes the SDK's hook-output JSON to stdout. Fail-open: any error -> emits `{}` (CLI's "no opinion") and exits 0, so a broken shim degrades to the CLI's normal behavior, never a hang/crash. - _hook_ipc.py: HookIpcServer (Unix domain socket preferred, TCP loopback fallback) that dispatches each event to the user's options.hooks callbacks (matching the old control-protocol hook input/output shapes, incl. async_/ continue_ conversion) and, for PreToolUse when can_use_tool is set, calls can_use_tool and translates PermissionResultAllow(updated_input=...)/ PermissionResultDeny(message=...) into the CLI's hookSpecificOutput.permissionDecision + updatedInput short-circuit. build_hooks_settings() synthesizes the settings.json hooks block wiring the relevant events to the shim (preserving user matchers). Token-authenticated; bounded reads; all dispatch failures return `{}`. Verified end-to-end (server + real shim subprocess): PreToolUse allow with updated_input rewrites the input, deny carries the reason, PostToolUse returns additionalContext, bad-token returns `{}`, user hooks fire with correct input shapes + tool_use_id. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Wire hooks + can_use_tool through the settings-hook IPC bridge (PTY) win #3 (part 2/3): the transport now starts the HookIpcServer in connect(), synthesizes a --settings hooks block wiring the relevant events to the shim, injects the shim's IPC endpoint into the child env, and tears the server down in close(). Programmatic hooks AND can_use_tool now run over this deterministic channel instead of (only) the TUI screen-scraper. pty_cli.py: - connect(): _start_hook_ipc() before _build_command(); env injection of CLAUDE_AGENT_SDK_HOOK_IPC in _build_env(); _inject_hook_settings() merges the synthesized hooks into an existing --settings JSON object (or appends one; non-object existing value -> second --settings, which the CLI merges). - _validate_options(): removed the `hooks` rejection (now supported). - _start_hook_ipc()/_normalized_hooks(): non-fatal startup; on failure the transport falls back to the existing watcher and hooks simply don't fire. - close(): stop the hook IPC server (cancel serve task + unlink socket; no leak). - Double-handling avoidance: _can_use_tool_via_hook flag makes _decide_permission allow-once WITHOUT re-invoking can_use_tool when a permission dialog reaches the watcher (the PreToolUse hook already decided server-side and the CLI short-circuits the dialog). _on_hook_permission_decision records a hook-channel deny into _pending_denied_tools so result.permission_denials still correlates. _hook_ipc.py: added the optional on_permission_decision sink (fires after a can_use_tool PreToolUse decision) so the transport can record hook-channel denies. Tests: - test_hook_ipc.py (new, 43): matcher semantics, settings synthesis, in-process dispatch (user hook shape, matcher filtering, async_/continue_ conversion, can_use_tool allow+updated_input / deny / raise->deny, decision sink, combined user-hook + can_use_tool), server lifecycle, real shim subprocess round-trip, bad-token and no-endpoint fail-open. - test_pty_integration.py::TestHookIpcBridge (new, 10, end-to-end vs a fake CLI that actually runs the shim): PreToolUse fires with correct input shape + tool_use_id; PreToolUse hook can deny (PostToolUse then skipped); can_use_tool deny via hook (callback gets full {file_path,content}+tool_use_id); can_use_tool allow with updated_input ACTUALLY rewrites the executed input (content -> REWRITTEN, PostToolUse then sees it); PostToolUse fires with tool_response. - test_pty_integration validation test flipped: hooks accepted (not rejected); permission_prompt_tool_name still rejected. - test_pty_transport: TestHookSettingsInjection (merge/append/non-object/no-bridge) + TestCanUseToolViaHookGuard (no double-invoke; deny recorded). Full suite 1220 passed / 5 skipped. ruff + mypy clean. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Document hooks + can_use_tool support; live A/B verification win #3 (part 3/3): update the PTY transport module docstring to reflect that programmatic hooks and can_use_tool are now supported via the settings-hook IPC bridge (was: "hooks unsupported / rejected"), and that the screen-scrape watcher is the fallback for when the bridge can't start (and always handles plan/AskUserQuestion/app dialogs, which are not PreToolUse hooks). Live A/B vs the old stream-json SDK (/tmp/venv_old), real `claude`: - updated_input ACTUALLY applied: a Write the model issued with content="ORIGINAL_MODEL_CONTENT" was rewritten by can_use_tool -> PermissionResultAllow(updated_input={content:"HOOK_REWROTE_THIS"}); the file on disk contained the REWRITTEN content (PostToolUse tool_response confirmed it). Baseline old SDK identical. - can_use_tool deny via the hook channel: tool not run, full input + real toolu_ id delivered, result.permission_denials carries the baseline shape. - PreToolUse hook input key-set is EXACTLY the baseline's [cwd, effort, hook_event_name, permission_mode, session_id, tool_input, tool_name, tool_use_id, transcript_path]; PostToolUse fires with tool_response. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix hook-IPC review findings W1-W4 (PTY transport) W1: gate the Unix-socket branch on hasattr(socket, "AF_UNIX") instead of the always-False hasattr(os, "AF_UNIX"), so POSIX uses a 0700-dir Unix domain socket (TCP loopback stays the fallback). W2: close the can_use_tool permission bypass. build_hooks_settings now ALWAYS emits a catch-all PreToolUse shim entry when can_use_tool is set (in addition to any narrow user matchers), so the callback is consulted for every tool. The IPC dispatch dedups can_use_tool per tool_use_id so a tool matching both the catch-all and the user's narrow matcher invokes the callback only once. W3: _close_stderr_pipe now closes both the write and read pipe ends (and clears _stderr_read_fd) on every failed-spawn path, so a direct caller that skips close() after a failed connect() does not leak the read fd. W4: _matching_callbacks fires non-tool-event hooks (UserPromptSubmit/Stop/...) regardless of the matcher pattern, mirroring the CLI's getMatchingHooks (undefined matchQuery -> all matchers). Settings synthesis already omits the tool matcher for non-tool events. Tests: +catch-all/dedup/unmatched-tool/no-double-invoke (W2), unix-socket+0700 (W1), failed-spawn fd cleanup (W3), non-tool-event matcher fires (W4). Gates green: ruff + mypy clean, pytest 1234 passed / 5 skipped. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH * Fix hook-IPC permission merge: deny-wins + per-id concurrency (W5/W6) W5 (MEDIUM): a can_use_tool ALLOW could silently overwrite a user PreToolUse hook DENY and leak a stale "hook-deny" reason into the allow (self-contradictory, under-enforced). _dispatch now routes the decision through HookIpcServer._merge_permission_decision: deny from EITHER source wins, and the permission triplet (permissionDecision / permissionDecisionReason / updatedInput) is replaced as an atomic unit so no stale field leaks. Non-permission hookSpecificOutput fields are kept. updated_input still applies on a both-allow result. W6 (LOW-MED): _run_can_use_tool_deduped held _cut_lock across the user can_use_tool callback, serializing distinct concurrent tool_use_ids (and risking deadlock on re-entrancy). Switched to double-checked locking with a per-tool_use_id _CutInflight marker: the lock now only guards the cache /in-flight dicts, distinct ids run the callback concurrently, and a duplicate fire for the same id awaits+replays the single result (exactly-once consult). perm is pre-bound so a raising callback cannot hang waiters. Bounded 512-entry FIFO cache retained. Tests: W5 deny-wins matrix (reason correctness + updated_input-on-allow + non-perm-field preservation); W6 distinct-id non-serialization, same-id single-consult, and raising-callback-wakes-waiters. https://claude.ai/code/session_01RVo5bksX7e8McFKqGbbuDH --------- Co-authored-by: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Modified the
ResultMessageclass to makecost_usdandtotal_cost_usdfields optional, preventing errors when API responses don't include cost information.Changes Made
cost_usdfield fromfloattofloat | None = NoneinResultMessageclasstotal_cost_usdfield fromfloattofloat | None = NoneinResultMessageclassProblem Solved
cost_usd,total_cost)client.py, which could cause initialization errors forResultMessageFiles Modified
claude_code_sdk/types.py: UpdatedResultMessagedataclass definitionclaude_code_sdk/_internal/client.py: Maintained consistency with existing commented-out cost fieldsBackward Compatibility
NonechecksNonewhen cost information is unavailableTesting
Related Issue
This change improves the robustness of the SDK when handling different API response formats and prevents potential runtime errors due to missing cost data.