Skip to content

feat(claude-producer): consume PostToolUse/Failure and PermissionDenied hooks (#76) - #173

Merged
tarakanof merged 4 commits into
mainfrom
feat/76-post-tool-and-permission-hooks
Sep 26, 2026
Merged

tarakanof merged 4 commits into
mainfrom
feat/76-post-tool-and-permission-hooks

Conversation

@tarakanof

@tarakanof tarakanof commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Closes #76

The owner chose to implement this directly instead of extending the spike. It keeps the issue's non-goal: no new producer states. The three events map onto running/waiting, or change only the trail.

Verified hook contract (2026-09-26)

Source: Claude Code hooks reference (fetched as hooks.md). The relevant sections are PostToolUse, PostToolUseFailure, PermissionDenied, PermissionRequest, matchers and async hooks.

  • All three events exist today. Each also matches on tool name, like PreToolUse. "", "*" or an omitted matcher matches everything. A matcher made of letters, digits, _, -, , and | is an exact name or a list of names. Anything else is an unanchored JS regex.
  • PostToolUse fires after a tool call succeeds. Fields: common ones plus tool_name, tool_input, tool_response (shape depends on the tool), tool_use_id, duration_ms.
  • PostToolUseFailure fires when a tool that started running fails. Fields: tool_name, tool_input, tool_use_id, error (string), is_interrupt (optional bool), duration_ms. For Bash, error starts with an Exit code N line, followed by the command's interleaved stdout and stderr. The docs say to key on tool_name, is_interrupt and that first line only. It does not fire for permission denials or validation rejections.
  • PermissionDenied fires only in auto mode when auto mode denies a call. It doesn't fire for a manual denial in the dialog, a PreToolUse-hook block or a deny rule. Fields: tool_name, tool_input, tool_use_id, reason (e.g. [Irreversible Local Destruction]). Its output: hookSpecificOutput.retry. We never emit that.
  • PermissionRequest has tool_name and tool_input but no tool_use_id. PreToolUse runs before every call, so the order is PreToolUse → PermissionRequest.
  • Timeouts and execution: command hooks default to a 600 s timeout, and matching hooks run in parallel. "async": true runs a command hook in the background: Claude doesn't wait, and timeout isn't enforced. Our own HookTimeoutMs still bounds the POST.
  • Spike data from the spike log (spike-hooks.jsonl) over about one day: 14,892 events, of which 14,474 were PostToolUse, 412 PostToolUseFailure (408 Bash) and 6 PermissionDenied. The error text contained full command output, so it must never be sent.

Design

  • The approved call's outcome ends waiting. Before this change, once a permission was approved nothing fired until the next hook. After the last tool call of a turn, the session kept showing blinking approve Bash until the next prompt. Now PermissionRequest records pending_permission in the marker: a sha256 of tool name + canonical tool_input. It's marker-only, never on the wire, and needed because there is no tool_use_id. When PostToolUse, PostToolUseFailure or PermissionDenied carries the same fingerprint while the session is waiting, the session goes back to running with one POST. A parallel call finishing (e.g. a Read) can't end another call's wait. The ~6 s permission_prompt Notification re-upserts waiting and keeps the fingerprint. Any other state clears it.
  • Failures and denials go in the trail, not a new state. The call's trail item gets a suffix: Bash: npm test (exit 1), (failed), (aborted) (when is_interrupt) or (denied). The new producer.AnnotateTrail finds the item by value (newest match) and replaces it in place, so an outcome that arrives late still marks the right item. The result is capped at 80 runes and rune-safe. In single-item mode (trail off), a late outcome never overwrites a newer call. With activity detail off, these hooks do nothing.
  • Attention semantics: a failed tool keeps the session running. Failing tests are routine, so a red ERR on every failed command would be noise; ERR stays reserved for StopFailure. An auto-denial doesn't blink: auto mode denied it and Claude continues, so there's nothing for the user to do. If a PermissionRequest did precede the denial, its wait ends.
  • Privacy: tool_response is never decoded (no struct field). From error, only the ^Exit code (\d{1,3})$ first line is read. The reason is not stored.
  • Server load and latency:
    • PostToolUse on a running session (about 97% of these events) does no marker write and no POST.
    • A failure or denial only rewrites the marker. The next POST carries it: the next hook, or the heartbeat within 10 s.
    • The only extra POST is the one that ends a wait, once per approved permission.
    • Measured in TestOutcomeHooks_RequestsPerToolCallUnchanged: 100 tool calls, with 10 failures and 1 denial, gave 100 POSTs before and 100 after.
    • All three hooks are now async: true, so they add no inline latency. They were synchronous during the spike. They don't depend on running in order: trail items are matched by value, and a wait ends only for its own call.
  • Spike removed: spike.go and the spike log writer are gone. configure/install and deconfigure/uninstall now delete the old spike log (spike-hooks.jsonl), because it holds full error output.

Review follow-ups (commit 37213dc)

  • Late prompt: the approved dialog's permission_prompt Notification can arrive after the resume. If it names the same tool (or has no message) within 15 s, it's dropped instead of re-sticking waiting. A new dialog always sends PermissionRequest first, which clears the resume record, so a real new prompt is never dropped.
  • tool_use_id: PreToolUse's tool_use_id and fingerprint are stored in the marker. A PermissionRequest whose fingerprint matches inherits that id. When both sides know the id, an outcome whose id differs can't end the wait. This covers an identical call retried while the first call's async outcome is still in flight.
  • Large stdin: hook stdin is stream-decoded with a 64 MiB cap. tool_response is skipped token by token and never stored. The old 1 MiB cap truncated big PostToolUse payloads (tool_use_id comes after tool_response), so those hooks silently did nothing.
  • Shared POST builder: a single wireRequest now builds the POST body for the hook, outcome and heartbeat paths.

Install / upgrade

  • The hook commands are unchanged from the spike, which shipped in v0.25.1 through v0.29.0. Installs from those versions start working as soon as the binary updates, with no settings change.
  • The next configure (CLI install/configure, or the app's Install toggle) rewrites our three entries in place with async: true. User hooks on the same events are kept, re-running changes nothing, and deconfigure removes all of ours. Covered by TestMergeSettings_OutcomeHooksAsyncAndUpgradeFromSpike.
  • The app's reconcile still doesn't re-run configure. I left that alone because nothing needs it: sync entries behave correctly, just without the latency win.
  • Installs older than v0.25.1 lack the three entries, so they behave as before until their next configure.

Tests

New posttool_test.go:

  • decoding of doc-shaped fixtures
  • the failureOutcome table, including the real "worktree isolated" refusal text → failed
  • rune-safe annotation
  • fingerprint key-order insensitivity
  • PostToolUse on a running session is a no-op; outcome hooks never create a session
  • approved → running, including after the permission_prompt Notification
  • a parallel call doesn't end the wait
  • a failure annotates the trail without a POST, and the heartbeat carries it
  • auto-mode denial; denial ending its own wait
  • detail off and trail off
  • privacy: no SECRET, sk-ant, error text or tool_response in any POST body (hook or heartbeat) or in the marker
  • the request-count comparison
  • install upgrade, idempotency and uninstall

Also new AnnotateTrail tests in internal/producer.

All tests use a temp HOME. None touches the real Claude settings file or launchctl.

gofmt -l (1.26) is clean, go vet ./... passes, and go test ./... -race passes.

Needs on-device verification

  • The fingerprint match between PermissionRequest and PostToolUse assumes both carry the same tool_input. If a PreToolUse hook rewrites the input (updatedInput), the fingerprints differ and the session falls back to today's behaviour (waits until the next hook). Worth checking in a real approve flow.
  • AskUserQuestion / ExitPlanMode (and any tool whose PostToolUse tool_input differs from its PermissionRequest's, e.g. answers or plan edits added): the fingerprint may not match, so the wait isn't ended early. The session then falls back to today's behaviour. Check which tools hit this in a real session.
  • A manual denial in the dialog fires no hook (per the docs), so that case still waits until the next hook, as it does today.

A tool's outcome can arrive after later calls were already prepended, so
the item is found by value (newest match) and replaced in place rather
than assumed to be the head. Groundwork for #76.
…rmissionDenied

Replaces the #76 log-only spike. A day of spike data (14,892 events: 97%
PostToolUse, 412 failures, 6 auto-mode denials) showed the value is not
in successes but in two gaps:

- An approved permission left the session in "waiting" ("approve Bash")
  until the next hook; after a turn's last tool call that was the next
  prompt. The approved call's PostToolUse(Failure), or its auto-denial,
  now ends the wait with one POST. The PermissionRequest records a hashed
  tool name + tool_input fingerprint in the marker so a parallel call's
  outcome can't end someone else's wait.
- Failures and denials were invisible. The call's trail item now gets a
  short suffix (exit N / failed / aborted / denied) in the marker only;
  the next POST carries it, so no extra requests (100 tool calls = 100
  POSTs before and after, see TestOutcomeHooks_RequestsPerToolCallUnchanged).

No new states: a failed tool stays running (ERR remains StopFailure's) and
a denial doesn't blink. tool_response is never decoded; the failure's
error text (the tool's output, per the spike log) is read only for its
'Exit code N' first line.

The three hooks are now registered async so Claude never waits on them;
an existing spike install is upgraded in place on the next configure.
Also corrects the stale Stop->DELETE mapping: Stop has been a no-op and
SessionEnd the delete since sessions stay until the window closes.

@tarakanof tarakanof left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: #173 (closes #76)

I read the hooks reference myself (hooks.md, fetched 2026-09-26) and checked each contract claim in the PR body. All of them match: the event names, the fields, tool_use_id missing from PermissionRequest, PostToolUseFailure not firing on permission denials, PermissionDenied firing only in auto mode, the Exit code N first-line guidance, and async: true meaning no wait and no enforced timeout.

Checks run: go test ./... -race with HOME set to a temp dir passes, and CI Go passes. The tests are isolated: hookHarness and the install tests use t.TempDir() plus t.Setenv("HOME"), and no test touches launchctl.

Verified OK

  • Async ordering. PermissionRequest is a synchronous decision hook, and it finishes before the dialog shows. PostToolUse fires only after the call has run. So an outcome can never be processed before its own PermissionRequest's marker write. Concurrent async hook processes are serialized by flock(LOCK_EX) on the per-session lock. enrichMarker (statusline) round-trips the marker struct, so it keeps pending_permission. The tick decodes into StatusRequest, so the field never goes out on the wire.
  • Privacy. There is no tool_response field, so it is never decoded. From error, only the anchored first line is read. reason is dropped. The fingerprint is an unsalted, truncated sha256. That can be guessed for short inputs, but it is marker-only (0600), and the same file already holds the plaintext Bash: <cmd> activity when detail is on, as do the transcripts under ~/.claude. Not a concern.
  • Load. TestOutcomeHooks_RequestsPerToolCallUnchanged gives 100/100. A PostToolUse on a running session does no write and no POST.
  • Install. Spike entries are upgraded in place, the user's PostToolUse hook is kept, re-running is byte-identical, and uninstall removes only ours.

Should-fix

  1. The spike log stays on disk with raw tool output. ~/.local/state/ember/spike-hooks.jsonl stored the full error field, which holds command stdout/stderr and may include secrets. It is 2.2 MB on the dev Mac. The PR deletes the writer but leaves the file, telling users to delete it by hand. Please os.Remove(spikeLogPath(home)) in configure/install and in uninstall. It costs one line and actually closes the privacy story.

Nits / known gaps (fallback is today's behaviour, so none is a regression)

  1. A late permission_prompt Notification re-sticks waiting. Verified with a scratch test: pre → permission-request → post-tool-use (resume) → notification gives state=waiting pending="". Later PostToolUse events no longer end it. It happens when the user approves right at about 6 s, a fast tool's async PostToolUse wins the lock, and the Notification (also fired in the background) lands afterwards. The window is narrow. A fix is to keep the resumed fingerprint as a short-lived tombstone and skip a permission_prompt that arrives while running just after a resume. At minimum, list this under on-device checks.
  2. An identical retried call can end the wrong wait. Verified: pre → PR(A) → pre → PR(B, same input) → the late failure of A gives running while B's dialog is up. It needs A's async outcome to land after B's PermissionRequest. That is rare, but likely when the POST is slow and holds the lock for up to HookTimeoutMs. Possible hardening: PreToolUse carries tool_use_id and runs synchronously just before PermissionRequest. Store it with the fingerprint and require the outcome's tool_use_id to match.
  3. Other fingerprint mismatches (not verified on device). AskUserQuestion's answers and an ExitPlanMode plan edited in the dialog may make PostToolUse's tool_input differ from PermissionRequest's. The session then waits until the next hook, as it does today. Add these to the "Needs on-device verification" list.
  4. The 1 MiB stdin cap drops some PostToolUse events. io.LimitReader(os.Stdin, 1<<20) truncates the input, and decoding then fails. Edit's tool_response includes originalFile, so an approved Edit on a file over about 1 MB never resumes. Rare.
  5. wireRequest duplicates the fresh-POST shaping in tick.go. The context_pct gate, the card/bar toggles and the weekly-field strip now live in two places. Extract a shared helper so the "weekly fields never on /v1/status" rule has one owner.

Also seen: gofmt -l with the local Go flags cmd/ember/device_display_test.go. That file is not in this diff and was already like that on main.

Verdict: it resolves #76 and I found no blockers. Fix #1 before merging (it's small). #2 to #6 can be follow-ups or doc notes.

- configure/deconfigure delete the #76 spike log: its error values hold
  full failed-command output.
- The approved dialog's late permission_prompt Notification (within 15 s
  of the resume, naming the same tool) no longer re-sticks waiting.
- PermissionRequest inherits the preceding PreToolUse's tool_use_id when
  fingerprints match; an outcome with a different id can't end the wait,
  so an identical call retried while the first's async outcome is in
  flight keeps its own wait.
- Hook stdin is stream-decoded with tool_response skipped token by token.
  The old 1 MiB cap cut off large PostToolUse payloads (tool_use_id comes
  after tool_response), so those hooks silently did nothing.
- One wireRequest POST-body builder shared by hook, outcome and heartbeat
  paths.
@tarakanof
tarakanof merged commit 3444da1 into main Sep 26, 2026
3 checks passed
@tarakanof
tarakanof deleted the feat/76-post-tool-and-permission-hooks branch September 26, 2026 19:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

producer(claude): evaluate PostToolUse + PermissionDenied hooks for richer activity

1 participant