Skip to content

Claude session idle-released while its run waits on a delegate_task child; Stop and Settle then fail permanently #15173

Description

@areidyOTH

What happened

A V2 Claude thread can't be stopped or settled. I press Stop and nothing happens; Settle is refused. Messages I send ("stop", "stop working") queue up and never start. The thread had used delegate_task to start a Codex reviewer, and that reviewer sat on an approval prompt.

Diagnosis

The Claude session is released after 30 minutes idle while its run is still waiting on an app-owned delegated child (delegate_task). After that release, nothing in the UI can end the run.

  1. The parent Claude run calls delegate_task (Codex child), and its native turn ends (end_turn). T3 correctly keeps the run running while it waits for the app-owned child. The child is slow; here it waited on an approval prompt.
  2. ProviderSessionManager.releaseIfStillIdle (apps/server/src/orchestration-v2/ProviderSessionManager.ts ~L864–940) only defers release for runtime.hasPendingBackgroundWork, the adapter's native background work. It doesn't count an app-owned delegated child of a running run. Exactly 30 minutes after end_turn the Claude session is released (query.close), and the session is persisted as stopped. The provider turn row stays running.
  3. Stop (run.interrupt) fails. In Orchestrator.ts the settleOnly fallback (~L7926) applies only when providerTurn.status !== "running" and the session is gone. A running turn with a released session falls through to the live-session path and fails with Provider session … is not active. (~L7978).
  4. thread.settle fails with has active or blocked work and cannot be settled because the run is still running.
  5. When the delegated child ends, its completion notice is queued as a new run behind the wedged run. User messages are queued the same way. Neither ever starts, and promote-to-steer fails for the same reason.

Only a server restart clears it (startup reconciliation ends unfinished runs). Restarting interrupts every other active thread on the server.

The same parent-session release probably explains the parent_not_active errors other threads got from delegate_task while their parent was waiting on a child.

Both code paths are unchanged on main as of 2026-10-03 (releaseIfStillIdle is identical; settleOnly still requires a non-running turn).

Steps to reproduce

  1. On a V2 Claude thread, have the agent call delegate_task for a Codex child that won't finish within 30 minutes. For example, use runtimeMode approval-required with a prompt that triggers a command approval, and leave the approval unanswered.
  2. Let the parent's Claude turn end while it waits for the child. Wait just over 30 minutes (default idle timeout). The provider event log shows an outgoing query.close, and the provider session becomes stopped. The run, run attempt, root node and provider turn all stay running.
  3. Press Stop on the parent thread. It fails with Provider session … is not active.
  4. Press Settle. It fails with has active or blocked work and cannot be settled.
  5. Send a message. It queues and never starts. Promoting it to steer fails.

Version

Desktop app 0.0.46-nightly.20261003.2610 (8ed276c); CLI 0.0.45

Environment

Linux x64 (kernel 7.0.0), desktop AppImage, Node v26.8.2; parent provider claudeAgent (claude-opus-5-5), delegated child codex

Evidence

# Provider event log (parent thread): turn ends, then idle release exactly 30 min later
[2026-10-03T10:31:30.363Z] incoming result stop_reason=end_turn
[2026-10-03T11:01:30.935Z] outgoing {"type":"query.close"}

# Projection state (statev2.sqlite) for the parent thread afterwards
run ordinal 1        status=running    (since 10:21:41Z, completed_at NULL)
provider_turn 1      status=running
provider_session     status=stopped    updated_at=11:01:30.936Z
subagent (codex, app_owned, delegate_task)  status=interrupted  completed_at=12:11:56Z
run ordinal 3 (delegated-completion notice)  status=queued
run ordinal 5 (user "stop working")          status=queued

# Server trace: 13 run.interrupt attempts between 12:07 and 12:14, all failed
OrchestratorDispatchError: Failed to dispatch orchestration command run.interrupt (...)
  [cause]: Error: Provider session provider-session:provider-instance:claudeAgent:thread:<id>:<uuid> is not active.

OrchestratorDispatchError: Failed to dispatch orchestration command thread.settle (...)
  cause: Thread <id> has active or blocked work and cannot be settled.

OrchestratorDispatchError: Failed to dispatch orchestration command queued-message.promote-to-steer (...)
  cause: Provider session ... is not active.

Related issues

No existing issue. Closed, unmerged PRs addressed parts of this: #14857 (Stop ends a V2 run whose session is gone, which is exactly the Stop half; closed 2026-10-02 without merging), #14507 (settle a run when its idle session is released; closed), and #14856 (event-consumer failure as a different trigger for the same stranded-run state). #14507's report had a different trigger (pinned native background work that expired after 4 h). This report adds a trigger that needs no failure at all: an app-owned delegate_task child that is still working when the parent's 30-minute idle release fires. #15124 is a different stuck-waiting-run cause (stalled checkpoint capture).

Fix applied or workaround

None applied yet. Workaround: delete the queued messages on the stuck thread so it doesn't resume unexpectedly, then restart the app (startup reconciliation ends the run), then settle the thread.

Filed by

Claude Code (Claude Opus 5.5) via t3 triage

Activity

  1. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @areidyOTH for the detailed diagnosis, the code pointers, and the timing evidence. I confirmed this on main (31a9da179). It's a real stuck state, and it's separate from #15124 (that one is a waiting run whose checkpoint capture never finishes).

    Why Stop, Settle, and steer fail

    This matches your report. dispatchRunInterrupt only takes the session-gone path when the provider turn is already not running (Orchestrator.ts, settleOnly). A turn that's still running after its session was released falls through and errors with Provider session … is not active. thread.settle then refuses the thread because a running run counts as active work. Promote-to-steer hits the same dead session, so the queued user message and the delegated-completion wake never start. Only a server restart recovers it, because startup reconciliation is what ends the unfinished run.

    Why the session was released

    Also as you described. releaseIfStillIdle defers only while hasPendingBackgroundWork is true, and for Claude that covers native background tasks, native session subagents, and wake buffers. An app-owned delegate_task child is none of those, so it doesn't pin the session, and the default 30-minute idle timer runs. The gap from end_turn at 10:31:30Z to query.close at 11:01:30Z matches that timer. The session pump marks the session idle when it sees turn.terminal, which starts the timer.

    One correction to the causal chain

    T3 doesn't intentionally keep the run and provider turn running after that terminal. A root turn.terminal that the run consumer ingests is persisted as waiting (RunExecutionService.writeFinalRunEvents), and the provider turn leaves running. Since your captured rows were still running, that terminal likely never landed on the projection, even though the session manager saw it and later released the session. That's the same stranded-projection class as the open PR #14856 and the closed, unmerged PRs #14507 (settle the run when the idle session is released) and #14857 (Stop settles a running turn whose session is gone). The app-owned child explains why the release wasn't deferred, but on its own it doesn't explain a provider turn that stays running after a persisted end_turn.

    The parent_not_active errors aren't established by this capture. delegate_task returns that when the caller has no active run owned by its provider, and this parent run was still running, which passes that check. Releasing the session revokes the MCP credential separately.

    Optional extra evidence

    If you still have the trace, whatever the server logged around 10:31Z when the end_turn result arrived would help narrow it down. An event-write failure would make this the #14856 trigger, while a live consumer that dropped the terminal would be a separate routing bug.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 3, 2026
  3. areidyOTH commented on Oct 3, 2026

    @areidyOTH
    Author

    Thanks. Correction on parent_not_active accepted.

    The server trace around 10:31Z has rotated out (only 3 spans survive for 10:21–10:23), but the persisted event log narrows it down. The run consumer stopped persisting at 10:22:21.623Z, nine minutes before end_turn, not at the terminal:

    • Last consumer-written events for the thread: seq 624482–624485, the first Bash tool call's node.updated/turn-item.updated (completed at 10:22:21.6Z). Its provider event log shows 53 tool_use blocks in total, assistant text, and the result at 10:31:30Z. None of the 52 later tool calls, the text, or the terminal were persisted.
    • The only later events on the thread are the delegate_task subagent/node/turn-item rows at 10:31:20.538Z (MCP command path), then provider-session.updated stopped at 11:01:30.936Z.
    • SQLite was writable in the gap: another thread persisted 53 events at 10:24:52–10:25:47Z, and this thread's MCP-path rows landed at 10:31:20Z. No other thread had provider activity in the gap.

    So it looks like the consumer stopped ingesting (or routing) mid-turn rather than the #14856 write-failure path, though a transient write failure at exactly 10:22:21 can't be excluded without the trace. One possibly relevant detail: a thread-send from another thread queued run 2 on the same provider thread at 10:21:50Z, while run 1 was still preparing (its provider turn started 10:22:09Z).

    Filed by Claude Code (Claude Opus 5.5) via t3 triage.

  4. postoso commented on Oct 5, 2026

    @postoso

    Another occurrence on macOS, 0.0.46-nightly.20261004.2657. One difference from the original report: neither delegated child was still pending when the session was released. Both had already completed, and the run stayed running anyway.

    Timeline (UTC, from the provider event log and statev2.sqlite):

    • 01:02:57 run ordinal 10 starts (claudeAgent, claude-opus-5-5)
    • 01:03:53 and 01:07:26 the agent starts two app-owned delegate_task children (codex)
    • 01:09:31 child 1 completes
    • 01:09:44.574 last provider-sourced turn item persisted for the run (a dynamic_tool)
    • 01:09:51 to 01:10:05 the provider log has one more thinking block, a tool_use and its tool_result, and the final assistant text. None of these are in orchestration_v2_projection_turn_items.
    • 01:10:06.611 result / success / terminal_reason: completed
    • 01:13:07 child 2 completes. Its subagent and notification turn items are persisted, so app-sourced writes for this run still worked.
    • 01:43:09.399 outgoing query.close, 30 minutes after the last child completion (not after end_turn)

    State afterwards: run, run attempt, and provider turn are all running; provider session stopped (updated 01:43:09.399Z); both subagent rows completed; no queued runs.

    run.interrupt and a restart-mode send (message.dispatch, via the t3_thread_send MCP tool) both fail the same way:

    OrchestratorDispatchError: Failed to dispatch orchestration command run.interrupt (...)
      [cause]: Error: Provider session provider-session:provider-instance:claudeAgent:thread:<id>:<uuid> is not active.
    

    On the triage question about an event-write failure vs. a dropped terminal: the provider-sourced events stopped persisting about 22s before the result, while app-sourced items kept landing three minutes later. That matches the consumer-stopped-early pattern in the comment above, and points at the #14856 class rather than the terminal alone being dropped. I can't tell you what the server logged at 01:09:44, because the trace had already rotated out. The 10 trace files cover about 15 minutes under this load.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions