Skip to content

[Bug]: Cancelling a delegated task leaves its Claude native subagents running forever (task stuck waiting_for_children) #17154

Description

@jnyross

Before submitting

Area

apps/server

Steps to reproduce

Summary

Cancelling an app-owned delegated task (delegate_task → task_cancel) whose Claude child turn has provider-native subagents in flight leaves those native subagent nodes running forever. The delegated task then stays status: running, workState: waiting_for_children indefinitely, and the sidebar keeps showing the child thread as Running with a growing timer (12h+ in the observed case). Settling and archiving the child thread, a second task_cancel, and t3_thread_interrupt (returns no_active_run) do not clear it.

Steps

  1. From a Claude-provider parent thread, delegate_task (mode async) to a claudeAgent child.
  2. Let the child launch several native Claude subagents (Agent/Task tool). In the observed case it launched 5 within ~70 s.
  3. While they are still running, task_cancel the delegated task.
  4. Observe: the child root turn ends (provider-turn.interrupt effect succeeds), but the native subagent nodes never receive a terminal status.

Observed state (read-only query of statev2.sqlite)

  • orchestration_v2_projection_subagents: 5 rows with origin: provider_native, status: running, completed_at: NULL, last updated_at ~2–5 min before the cancel. The app-owned row for the delegated task itself is also still running.
  • orchestration_v2_projection_nodes for the child thread: 150 tool_call completed, 7 assistant_message completed, 5 root_turn cancelled, 1 root_turn completed, 5 subagent running.
  • task_status: status: running, workState: waiting_for_children, hasPendingChildRuns: false, latestTerminalStatus: completed.
  • t3_thread_list (statuses running/waiting) does not list the child thread, yet the UI shows it Running.

Expected behavior

Cancelling/interrupting a delegated task (or its child root turn) should mark that turn's in-flight provider-native subagent nodes terminal (cancelled/interrupted), so the task reaches a terminal state and the sidebar stops showing it as Running. Failing that, there should be a supported way to stop/close stale subagent nodes.

Actual behavior

Native subagent nodes stay running with no reconciler; the parent task waits for them forever; no tool can clear it.

Runtime or environment

  • t3 v0.0.46-nightly.20261003.2638, Linux, v2 orchestrator
  • Parent and child provider: claudeAgent (child model claude-fable-5-1); cancel issued via the t3-code MCP task_cancel

Workaround

None found without editing the state database. Restarting T3 not yet tried.

Activity

  1. juliusmarminge commented on Oct 8, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the detailed report and the projection dumps. I checked this against main at 37eaf5d. I haven't reproduced it locally, but the code has a gap that fits what you saw.

    What task_cancel does

    • task_cancel dispatches thread.stop on the child, then stopDelegatedTasks for the app-owned tasks under it (OrchestratorMcpService.ts#L2005-L2034). A task in waiting_for_children always passes the "cancellable" check (#L2070-L2078), so a second cancel is accepted but has nothing to act on.
    • thread.stop interrupts a running run. If no run is running, it targets the latest run only when derivePendingBackgroundWork still lists something (Orchestrator.ts#L9034-L9048). Otherwise it just records the stop and holds the thread (#L9080-L9092).
    • With a live session, the interrupt only queues provider-turn.interrupt (Orchestrator.ts#L8991-L9006). After the adapter call returns, the worker dispatches thread.background-work.settle (EffectWorker.ts#L165-L195).

    Where the native subagents likely fall through (from reading the code)

    • dispatchBackgroundWorkSettle calls settleInterruptedRun only when the stopped run is still running (Orchestrator.ts#L8584-L8599). settleInterruptedRun is the path that cascades a terminal status to run-owned native subagent rows and nodes (#L8279-L8317), and even then only for subagents whose runId is that run.
    • Otherwise it falls through to settleBackgroundWork. That marks pending background turn items interrupted (including subagent items) and clears the roster (Orchestrator.ts#L8134-L8191). It emits no subagent.updated or node.updated, so the linked orchestration_v2_projection_subagents row and the subagent node stay running.
    • The live-turn interrupt cascade in RunExecutionService likewise covers only subagents reported with runId === run.id while that run's ingestion is still open (RunExecutionService.ts#L629-L638, #L1041-L1056).
    • The Claude adapter doesn't fill that gap. Stop on a settled turn closes the CLI process and resets the roster (ClaudeAdapterV2.ts#L7652-L7678), and on query exit it ends the subagents' open tool calls (#L4953-L4968, #L7434-L7442). It never emits a terminal subagent.updated for the tasks in sessionSubagentsByTaskId, so they also stay running in the session registry (#L7845-L7849).

    Startup recovery already handles this case. Its comment says "Cancelling only the turn item would leave the linked subagent entity non-terminal forever", and it cancels the linked subagent row and node along with the item (ProviderRuntimeRecoveryService.ts#L489-L529). The Stop settle path has no equivalent.

    Why it then stays stuck

    • delegatedTaskProgress returns waiting_for_children whenever any subagent row on the child is active, whatever its origin (SubagentProjection.ts#L237-L256), and task_status reads the child's subagent rows directly (OrchestratorMcpService.ts#L1234-L1243). That matches your running / waiting_for_children / hasPendingChildRuns: false / latestTerminalStatus: completed output.
    • Stop, the thread list and t3_thread_interrupt look at runs, turn items and the roster instead. derivePendingBackgroundWork explicitly "does not consult subagent entities" (orchestrationV2PendingBackgroundWork.ts#L186-L222). Once the turn items are interrupted there is nothing for a later Stop to target, which likely explains no_active_run and a second task_cancel doing nothing.

    My best guess at your sequence: the child turn had already settled with the native agents still running in the background, so Stop took the settled-run path above. I couldn't confirm that from the dump alone. If it holds, a restart may not clear it either, because recovery's stale sweep keys off non-terminal turn items, and those were already marked interrupted.

    Related

    Suggested fix (untested)

    In settleBackgroundWork (Orchestrator.ts#L8134-L8164), when a subagent turn item is interrupted, also emit subagent.updated (status interrupted) for the linked non-app-owned row and node.updated for its node, the way startup recovery does. Do the same for open provider_native rows on the stopped thread whose turn item is already terminal, so state that's already stuck gets repaired by the next Stop. #16814's adapter-side cleanup could then also run when Stop closes the process (ClaudeAdapterV2.ts#L7665-L7677), not only at the next turn start. A regression test should cover a Claude child that settles with background native agents and is then cancelled via task_cancel, asserting the subagent rows and nodes end terminal and the task reaches a terminal state.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions