Repository navigation
fix(server): Stop ends a V2 run whose agent session is already gone - #14857
Closed
SunkenInTime wants to merge 2 commits into
Closed
SunkenInTime wants to merge 2 commits into
SunkenInTime wants to merge 2 commits into
Conversation
A run can stay "running" after its provider session is released, for example when its event consumer died. Stop then failed with "Provider session ... is not active", so the user could not end a turn that had no process behind it, and only a server restart cleared the thread. Stop now settles that run as interrupted directly: its provider turn, attempt, open nodes, streaming messages and background work. A session that is gone has nothing left to stop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is a contained server-side fix for an existing Stop path that previously left orphaned runs stuck, with targeted coverage for run, child-work, streaming-item, and checkpoint cleanup. Normal live-session behavior and product defaults remain unchanged, and no sensitive, schema, infrastructure, or static-analysis configuration is affected. You can add or adjust custom eligibility rules. Learn more. |
…agent threads Stop on a run whose session is gone left the run's assistant and reasoning items running, because the Stop projection only loads background-capable items. It also left a provider-native subagent's own thread with running nodes and items. Close the run's other open items, and terminalize provider-native subagents through the same cascade normal run finalization uses, including their child threads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When a V2 run is stuck as
runningand its provider session is gone, pressing Stop fails with "Provider session ... is not active". The thread keeps saying "Working". There's no way out short of restarting the server.This happens after a run loses its event consumer (see #14856). Nothing ever reports the turn ending, so the session manager eventually releases the idle session underneath it.
dispatchRunInterruptalready settles a run directly when its turn is not running and the session is gone. But a turn that is still markedrunningfalls through to the live-session path and errors out.flowchart TD S[User presses Stop] --> T{provider turn running?} T -- no --> G1{session gone?} G1 -- yes --> SO[settle background work only] G1 -- no --> I[send provider interrupt] T -- yes --> G2{session gone?} G2 -- no --> I G2 -- "yes (before)" --> E["error: Provider session ... is not active"] G2 -- "yes (after)" --> X[settle the run as interrupted]Change
If Stop targets a running turn whose provider session is gone, there is no process left to interrupt, so Stop settles the run itself:
interruptedcascadeTerminalizeRunOwnedSubagentsthat normal run finalization uses, because they died with the provider process. App-owned delegated children run in their own threads and are left alone, since they may still be working.checkpoint.captureis enqueued, the same rollback point a normal interrupt recordsThe existing settle-only branch is unchanged. It now shares one session-liveness check with the new branch.
Relation to #14856 and #14365
These are three independent layers, and each can merge on its own:
runningwith no session behind it, whatever the cause.No files overlap with #14365.
Scope and approval
No issue yet. This is a focused fix. Stop errored on a state the server already treats as dead in the neighbouring branch, and a maintainer hit it on a real thread.
Verification
Before, on the current Alpha build. The run's consumer died, the session idled out, and Stop errors while the thread keeps "Working":
After, on this branch. I used the same setup, a 30s SQLite lock mid-turn so the consumer dies, then waited for the session to idle out. For this recording only, I set the idle timeout to 60s locally instead of 30 minutes and did not commit that. The session showed
stoppedat 18:39:19 with the run stillrunning. Pressing Stop settled it at 18:40:14.779.Recording of the Stop press:
https://files.daracloud.uk/files/t3code-v2-lost-runs/26be299e19ed-pr2-after-stop.mp4
Database after Stop: run
interrupted, provider turninterrupted,run_interrupt_resultwith the message above.Tests in
Orchestrator.control-reads.test.ts:stops a running turn whose provider session is already gone. It seeds a running run, attempt, root node with a checkpoint scope, provider turn, running command item, streaming assistant item, one provider-native subagent whose child thread has a running node and item, one app-owned subagent, and a provider thread pointing at a session the server does not hold. It asserts all root records, both items, the native subagent and its child thread's node and item endinterrupted, the app-owned child is stillrunning, and onecheckpoint.captureeffect exists. On the old code it fails with "Provider session session:orphaned-run:released is not active."All test files that dispatch
run.interruptpass (193 tests). Every interrupt and Stop test undersrc/orchestration-v2passes (147 tests). Scoped lint and servertsc --noEmitare clean for the changed files.Known gap: pending runtime requests (approvals) are not part of this projection query. Session release already expires them, so one would only survive if that release write also failed.
Investigation, fix and verification were done by Claude Opus 5.5 in Claude Code inside T3 Code. Independent review was done by GPT-6 Astra through Codex. I added subagent and checkpoint handling from its findings before opening. Macroscope then caught open assistant items and native child threads, fixed in e5c15c7.
🤖 Generated with Claude Code