Skip to content

fix(server): Stop ends a V2 run whose agent session is already gone - #14857

Closed
SunkenInTime wants to merge 2 commits into
pingdotgg:t3code/codex-turn-mappingfrom
SunkenInTime:fix/v2-stop-settles-orphaned-runs
Closed

SunkenInTime wants to merge 2 commits into
pingdotgg:t3code/codex-turn-mappingfrom
SunkenInTime:fix/v2-stop-settles-orphaned-runs

Conversation

@SunkenInTime

@SunkenInTime SunkenInTime commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Problem

When a V2 run is stuck as running and its provider session is gone, pressing Stop fails with "Provider session ... is not active". The thread keeps saying "Working". There's no way out short of restarting the server.

This happens after a run loses its event consumer (see #14856). Nothing ever reports the turn ending, so the session manager eventually releases the idle session underneath it. dispatchRunInterrupt already settles a run directly when its turn is not running and the session is gone. But a turn that is still marked running falls through to the live-session path and errors out.

flowchart TD
    S[User presses Stop] --> T{provider turn running?}
    T -- no --> G1{session gone?}
    G1 -- yes --> SO[settle background work only]
    G1 -- no --> I[send provider interrupt]
    T -- yes --> G2{session gone?}
    G2 -- no --> I
    G2 -- "yes (before)" --> E["error: Provider session ... is not active"]
    G2 -- "yes (after)" --> X[settle the run as interrupted]
Loading

Change

If Stop targets a running turn whose provider session is gone, there is no process left to interrupt, so Stop settles the run itself:

  • interrupt request and result items (the result says "Stopped. The agent's session had already ended.")
  • the provider turn, the run attempt and the run become interrupted
  • open nodes, streaming messages, and the run's other open items (assistant text, reasoning) are closed
  • provider-native subagents and their own child threads go through the same cascadeTerminalizeRunOwnedSubagents that normal run finalization uses, because they died with the provider process. App-owned delegated children run in their own threads and are left alone, since they may still be working.
  • the stopped run's checkpoint.capture is enqueued, the same rollback point a normal interrupt records
  • background work and the completion cohort settle through the existing helpers

The existing settle-only branch is unchanged. It now shares one session-liveness check with the new branch.

Relation to #14856 and #14365

These are three independent layers, and each can merge on its own:

No files overlap with #14365.

Scope and approval

No issue yet. This is a focused fix. Stop errored on a state the server already treats as dead in the neighbouring branch, and a maintainer hit it on a real thread.

Verification

Before, on the current Alpha build. The run's consumer died, the session idled out, and Stop errors while the thread keeps "Working":

Stop shows Provider session is not active and the thread keeps working

After, on this branch. I used the same setup, a 30s SQLite lock mid-turn so the consumer dies, then waited for the session to idle out. For this recording only, I set the idle timeout to 60s locally instead of 30 minutes and did not commit that. The session showed stopped at 18:39:19 with the run still running. Pressing Stop settled it at 18:40:14.779.

Run interrupted, Stopped. The agent's session had already ended.

Recording of the Stop press:

https://files.daracloud.uk/files/t3code-v2-lost-runs/26be299e19ed-pr2-after-stop.mp4

Database after Stop: run interrupted, provider turn interrupted, run_interrupt_result with the message above.

Tests in Orchestrator.control-reads.test.ts:

  • stops a running turn whose provider session is already gone. It seeds a running run, attempt, root node with a checkpoint scope, provider turn, running command item, streaming assistant item, one provider-native subagent whose child thread has a running node and item, one app-owned subagent, and a provider thread pointing at a session the server does not hold. It asserts all root records, both items, the native subagent and its child thread's node and item end interrupted, the app-owned child is still running, and one checkpoint.capture effect exists. On the old code it fails with "Provider session session:orphaned-run:released is not active."

All test files that dispatch run.interrupt pass (193 tests). Every interrupt and Stop test under src/orchestration-v2 passes (147 tests). Scoped lint and server tsc --noEmit are clean for the changed files.

Known gap: pending runtime requests (approvals) are not part of this projection query. Session release already expires them, so one would only survive if that release write also failed.

Investigation, fix and verification were done by Claude Opus 5.5 in Claude Code inside T3 Code. Independent review was done by GPT-6 Astra through Codex. I added subagent and checkpoint handling from its findings before opening. Macroscope then caught open assistant items and native child threads, fixed in e5c15c7.

🤖 Generated with Claude Code

A run can stay "running" after its provider session is released, for
example when its event consumer died. Stop then failed with "Provider
session ... is not active", so the user could not end a turn that had no
process behind it, and only a server restart cleared the thread.

Stop now settles that run as interrupted directly: its provider turn,
attempt, open nodes, streaming messages and background work. A session
that is gone has nothing left to stop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Oct 2, 2026
Comment thread apps/server/src/orchestration-v2/Orchestrator.ts
Comment thread apps/server/src/orchestration-v2/Orchestrator.ts Outdated
@macroscopeapp

macroscopeapp Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at e5c15c7

Macroscope's review found this PR approvable — This is a contained server-side fix for an existing Stop path that previously left orphaned runs stuck, with targeted coverage for run, child-work, streaming-item, and checkpoint cleanup. Normal live-session behavior and product defaults remain unchanged, and no sensitive, schema, infrastructure, or static-analysis configuration is affected.

You can add or adjust custom eligibility rules. Learn more.

@juliusmarminge juliusmarminge added the macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews label Oct 2, 2026 — with ChatGPT Codex Connector
…agent threads

Stop on a run whose session is gone left the run's assistant and
reasoning items running, because the Stop projection only loads
background-capable items. It also left a provider-native subagent's own
thread with running nodes and items. Close the run's other open items,
and terminalize provider-native subagents through the same cascade normal
run finalization uses, including their child threads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews size:L 100-499 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants