Skip to content

[Bug]: Claude child task dies after t3_worktree_handoff; continuation turn never starts #15136

Description

@daniifrim

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/desktop

Steps to reproduce

  1. In a Git project, from a parent thread, call delegate_task (mode async) with target claudeAgent / claude-opus-5-5 (also reproduced with claude-sonnet-5-5).
  2. The child's task: call t3_worktree_status, then, as its last action, t3_worktree_handoff with branch: feature/x, baseRef: main, startFromOrigin: false, runSetupScript: false and a continuationPrompt (for example "run pwd and report").
  3. Wait for the child.

Expected behavior

As with a Codex child (same prompt, codex / gpt-6.1-sol): the worktree is created, run 2 (the continuation) starts inside the worktree, and the parent gets one "Delegated task … reached a terminal state" notice after run 2.

Actual behavior

  • The worktree and branch are created.
  • About 2 seconds after the handoff tool call, T3 sends query.close to the Claude session. Run 1 then ends failed with: "The provider event stream closed unexpectedly. Retry the turn; if it keeps failing, check the provider and server logs." The handoff tool item ends cancelled.
  • Run 2 (the continuation) stays queued and never starts (startedAt null). On build 2610 it was once cancelled instead.
  • The parent is never woken, or is woken with status: failed. The task appears to be working forever.

It looks like the planned session detach on a workspace change is treated as an unexpected stream close for the Claude driver, which then blocks the queued continuation. The Codex driver handles the same detach correctly.

Impact

Major degradation or frequent failure

Version or commit

T3 Code 0.0.46-nightly.20261003.2610 and 0.0.46-nightly.20261003.2623 (3 of 3 Claude runs failed; 1 of 1 Codex run succeeded).

Environment

Linux (Debian 13, kernel 6.12), t3 serve web mode. Claude Code 2.1.288, codex-cli 0.159.1. Runtime mode full-access.

Logs or stack traces

Provider event log for the child (`logs/provider/events.thread-delegated-task-…log`), last events:

09:52:45.062Z incoming assistant: tool_use mcp__t3-code__t3_worktree_handoff
09:52:45.076Z incoming stream_event: message_delta stop_reason=tool_use
09:52:45.076Z incoming stream_event: message_stop
09:52:47.419Z outgoing query.close

Turn items: `dynamic_tool mcp__t3-code__t3_worktree_handoff` → `cancelled`; `error` → `failure.message: "The provider event stream closed unexpectedly. …"`, `class: unknown`. Runs: `ordinal 1 failed`, `ordinal 2 queued`.

Screenshots, recordings, or supporting files

No response

Workaround

Launch Claude workers with t3_thread_launch and workspaceStrategy: {type: "worktree", …} instead of delegate_task + handoff.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 3, 2026
  2. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @daniifrim for the careful Claude vs Codex comparison and the provider event log! This is a real bug on current main, and the split you saw matches what's in the code. I didn't find an existing issue for it, and the 0.0.46 nightlies you're on aren't behind a fix.

    What happens

    t3_worktree_handoff rebinds the thread and then deliberately queues the continuationPrompt. The metadata update schedules provider-session.detach with reason "Workspace changed." That detach is meant to end the current turn, and the queued run is meant to start the next one inside the worktree.

    Claude sessions are exclusive (supportsMultipleProviderThreadsPerSession: false), so for them detach releases the whole session. Releasing it fails every event subscriber with ProviderAdapterEventStreamError, and the Claude session finalizer sends query.close. That's the outgoing query.close you saw about two seconds after the handoff tool_use. The turn is then stored as failed with class unknown and "The provider event stream closed unexpectedly…". The handoff tool item is cancelled because the process is gone before its result can be delivered.

    startNextQueuedRun treats any non-validation failure on the same provider as a sign the next message will fail too, so it sets queueHeld on the continuation. That run stays queued with startedAt null until someone explicitly resumes the queue. A restart won't start it either.

    While the continuation is still queued, the child counts as "working", so the parent never wakes. If the detach settles run 1 before the continuation row exists, the parent wakes with failed instead. Both match what you saw.

    Codex doesn't hit this because its sessions allow multiple threads. There, detach interrupts the active turn and unloads that thread instead of failing the session stream. An interrupted turn doesn't hold the queue, so run 2 starts. The same exclusive-session release also applies to Cursor, OpenCode, Pi, and ACP, but not to Codex or OpenCode 2.

    Likely fix area

    Holding the queue makes sense after a real provider failure. A planned workspace detach could instead settle the active turn as an intentional stop, so the handoff continuation still gets promoted. A maintainer will decide on the fix direction.

    Related: #15135 (open) is the same "stream closed unexpectedly" outcome from a planned detach on an exclusive session, but it's triggered by an agent archiving its own thread.

    Workaround

    Your workaround is the right one: launch the worker with t3_thread_launch and a worktree workspaceStrategy, so the session starts inside the worktree. For a child that's already stuck, the thread's resume control should start the held continuation.

  3. added
    via-triageFiled through npx t3 triage
    and removed
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 3, 2026
  4. mmmattG commented on Oct 4, 2026

    @mmmattG

    Confirmed this also affects a normal top-level Claude thread; delegated child tasks are not required. The thread had parentThreadId: null and relationshipToParent: null.

    On 0.0.46-nightly.20261004.2644, using claudeAgent / claude-opus-5-5 with full access, the agent called t3_worktree_handoff with a continuation prompt. The worktree was created and the thread binding updated, but the original run failed with “The provider event stream closed unexpectedly.” The continuation remained queued with queueHeld: true and startedAt: null.

    This confirms the impact extends to ordinary conversations that move into a worktree mid-thread, beyond the delegated-task reproduction in this issue.

    — Matt’s agent, GPT-6.1 Sol (xhigh reasoning), running in T3 Code through the Codex harness, posting on behalf of @mmmattG.

  5. FranciscoJSBarragan commented on Oct 4, 2026

    @FranciscoJSBarragan

    Still reproduces on 0.0.46-nightly.20261004.2657 (macOS, desktop app), on a top-level thread with claudeAgent / claude-sonnet-5-5 (effort low) in full-access.

    The agent called t3_worktree_handoff (startFromOrigin: false, runSetupScript: false, explicit path, with a continuationPrompt). The worktree and the thread binding were created. Run 1 then ended failed ("The provider event stream closed unexpectedly"). Run 2 (the continuation) stayed queued with startedAt: null for more than 10 minutes, and t3_queue_list still showed it. An earlier 0.0.46 nightly, from before updating, behaved the same way.

    Sending a new message with t3_thread_send (mode: auto) starts a fresh run, but the held continuation stays queued and orphaned. It has to be cancelled with t3_queue_cancel, or it could run later as a duplicate.

    Our workaround: never pass continuationPrompt to the handoff, and prefer starting the thread already bound with t3_thread_launch + workspaceStrategy. That path works as expected (delegate_task children run in the bound checkout).

  6. SaulMoro commented on Oct 6, 2026

    @SaulMoro

    Still reproduces on 0.0.46-nightly.20261005.2702 (macOS desktop, top-level thread, claudeAgent / claude-opus-5-5, full-access):
    run 8 ended failed ("provider event stream closed unexpectedly") ~7 s after the t3_worktree_handoff tool_use,
    right after provider-session.detached reason "Workspace changed.". Difference from earlier reports: the worktree-continuation
    run started ~6 s later in the new cwd without being held, so on this build only the failed-turn half reproduced here.
    8/8 agent-initiated handoffs on this machine since 2026-10-05 21:41Z show the failed turn.

  7. JonathanEspinosaLong commented on Oct 8, 2026

    @JonathanEspinosaLong

    Still reproduces on 0.0.46-nightly.20261008.2819 (macOS 15.6 desktop, top-level thread, claudeAgent / claude-opus-5-5, full-access, Claude Code 2.1.294). This is the third agent-initiated handoff on this machine since 2026-10-07, and all three failed.

    t3_worktree_handoff with startFromOrigin: true, runSetupScript: false, explicit path, and a continuationPrompt. The provider log shows the handoff tool_use and message_stop at 14:56:15.8Z, then outgoing query.close at 14:56:20.7Z. Run 3 ended failed ("The provider event stream closed unexpectedly"). The continuation, run 4, is still queued with startedAt: null, and t3_queue_list still lists it. A later user message started run 5 ahead of it. The handoff dynamic_tool item is still running, which matches #15993.

    The worktree and thread binding were created correctly, so only the auto-resume is broken.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions