Repository navigation
[Bug]: Cancelling a delegated task leaves its Claude native subagents running forever (task stuck waiting_for_children) #17154
Description
Activity
Note
Grok responding on behalf of Julius.
Thanks for the detailed report and the projection dumps. I checked this against
mainat37eaf5d. I haven't reproduced it locally, but the code has a gap that fits what you saw.What
task_canceldoestask_canceldispatchesthread.stopon the child, thenstopDelegatedTasksfor the app-owned tasks under it (OrchestratorMcpService.ts#L2005-L2034). A task inwaiting_for_childrenalways passes the "cancellable" check (#L2070-L2078), so a second cancel is accepted but has nothing to act on.thread.stopinterrupts a running run. If no run is running, it targets the latest run only whenderivePendingBackgroundWorkstill lists something (Orchestrator.ts#L9034-L9048). Otherwise it just records the stop and holds the thread (#L9080-L9092).- With a live session, the interrupt only queues
provider-turn.interrupt(Orchestrator.ts#L8991-L9006). After the adapter call returns, the worker dispatchesthread.background-work.settle(EffectWorker.ts#L165-L195).
Where the native subagents likely fall through (from reading the code)
dispatchBackgroundWorkSettlecallssettleInterruptedRunonly when the stopped run is stillrunning(Orchestrator.ts#L8584-L8599).settleInterruptedRunis the path that cascades a terminal status to run-owned native subagent rows and nodes (#L8279-L8317), and even then only for subagents whoserunIdis that run.- Otherwise it falls through to
settleBackgroundWork. That marks pending background turn itemsinterrupted(includingsubagentitems) and clears the roster (Orchestrator.ts#L8134-L8191). It emits nosubagent.updatedornode.updated, so the linkedorchestration_v2_projection_subagentsrow and the subagent node stayrunning. - The live-turn interrupt cascade in
RunExecutionServicelikewise covers only subagents reported withrunId === run.idwhile that run's ingestion is still open (RunExecutionService.ts#L629-L638,#L1041-L1056). - The Claude adapter doesn't fill that gap. Stop on a settled turn closes the CLI process and resets the roster (
ClaudeAdapterV2.ts#L7652-L7678), and on query exit it ends the subagents' open tool calls (#L4953-L4968,#L7434-L7442). It never emits a terminalsubagent.updatedfor the tasks insessionSubagentsByTaskId, so they also stayrunningin the session registry (#L7845-L7849).
Startup recovery already handles this case. Its comment says "Cancelling only the turn item would leave the linked subagent entity non-terminal forever", and it cancels the linked subagent row and node along with the item (
ProviderRuntimeRecoveryService.ts#L489-L529). The Stop settle path has no equivalent.Why it then stays stuck
delegatedTaskProgressreturnswaiting_for_childrenwhenever any subagent row on the child is active, whatever its origin (SubagentProjection.ts#L237-L256), andtask_statusreads the child's subagent rows directly (OrchestratorMcpService.ts#L1234-L1243). That matches yourrunning/waiting_for_children/hasPendingChildRuns: false/latestTerminalStatus: completedoutput.- Stop, the thread list and
t3_thread_interruptlook at runs, turn items and the roster instead.derivePendingBackgroundWorkexplicitly "does not consult subagent entities" (orchestrationV2PendingBackgroundWork.ts#L186-L222). Once the turn items are interrupted there is nothing for a later Stop to target, which likely explainsno_active_runand a secondtask_canceldoing nothing.
My best guess at your sequence: the child turn had already settled with the native agents still running in the background, so Stop took the settled-run path above. I couldn't confirm that from the dump alone. If it holds, a restart may not clear it either, because recovery's stale sweep keys off non-terminal turn items, and those were already marked
interrupted.Related
- fix(server): Claude subagents from a closed process no longer stay running #16814 (open) is the closest fix. It makes the Claude adapter interrupt subagents left by an earlier, closed process, but only when the next turn starts on a new process. A stopped delegated child is held and never gets a next turn, so on its own it likely wouldn't clear this case.
- [Bug]: Cancelling a background-only delegated task reports completed with an interim summary #16739: the same "provider turn ended, background work pending, then cancel" shape, with a different symptom (the task reports
completed). - [Bug]: Nested Claude subagent stays Running after the SDK reports it completed during a user turn #16355 (lost nested completion), [Bug]: Provider turn projection stays running after T3 ends the turn #16557 (provider turn projection), Delegated task reports finished while its child thread still has running work #16603 (task finishes too early) and [Bug]: Claude background subagents stay Running forever when their wake continuation runs on another provider #17099 (subagents stuck when the wake runs on another provider) are nearby but have different causes.
- fix(orchestration): reconcile stale delegated task and held queue state #15201 changes
delegatedTaskProgressfor held queues but not how native subagent rows are settled.
Suggested fix (untested)
In
settleBackgroundWork(Orchestrator.ts#L8134-L8164), when asubagentturn item is interrupted, also emitsubagent.updated(statusinterrupted) for the linked non-app-owned row andnode.updatedfor its node, the way startup recovery does. Do the same for openprovider_nativerows on the stopped thread whose turn item is already terminal, so state that's already stuck gets repaired by the next Stop. #16814's adapter-side cleanup could then also run when Stop closes the process (ClaudeAdapterV2.ts#L7665-L7677), not only at the next turn start. A regression test should cover a Claude child that settles with background native agents and is then cancelled viatask_cancel, asserting the subagent rows and nodes end terminal and the task reaches a terminal state.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 8, 2026 - added a commit that references this issue
on Oct 8, 2026
Before submitting
Area
apps/server
Steps to reproduce
Summary
Cancelling an app-owned delegated task (
delegate_task→task_cancel) whose Claude child turn has provider-native subagents in flight leaves those native subagent nodesrunningforever. The delegated task then staysstatus: running,workState: waiting_for_childrenindefinitely, and the sidebar keeps showing the child thread as Running with a growing timer (12h+ in the observed case). Settling and archiving the child thread, a secondtask_cancel, andt3_thread_interrupt(returnsno_active_run) do not clear it.Steps
delegate_task(mode async) to aclaudeAgentchild.task_cancelthe delegated task.provider-turn.interrupteffect succeeds), but the native subagent nodes never receive a terminal status.Observed state (read-only query of statev2.sqlite)
orchestration_v2_projection_subagents: 5 rows withorigin: provider_native,status: running,completed_at: NULL, lastupdated_at~2–5 min before the cancel. The app-owned row for the delegated task itself is also stillrunning.orchestration_v2_projection_nodesfor the child thread: 150 tool_call completed, 7 assistant_message completed, 5 root_turn cancelled, 1 root_turn completed, 5 subagent running.task_status:status: running,workState: waiting_for_children,hasPendingChildRuns: false,latestTerminalStatus: completed.t3_thread_list(statuses running/waiting) does not list the child thread, yet the UI shows it Running.Expected behavior
Cancelling/interrupting a delegated task (or its child root turn) should mark that turn's in-flight provider-native subagent nodes terminal (cancelled/interrupted), so the task reaches a terminal state and the sidebar stops showing it as Running. Failing that, there should be a supported way to stop/close stale subagent nodes.
Actual behavior
Native subagent nodes stay
runningwith no reconciler; the parent task waits for them forever; no tool can clear it.Runtime or environment
task_cancelWorkaround
None found without editing the state database. Restarting T3 not yet tried.