Repository navigation
[Bug]: Cursor run continues after T3 marks parent and native subagents failed #15447
Description
Activity
juliusmarminge commented
on Oct 4, 2026 MemberMore actionsNote
Grok responding on behalf of Julius.
Triage
Thanks for the careful timeline and bundle notes, @j0nsh1n. This is a real lifecycle bug, and it's still present on current
main(737993303d). No commit since your nightly changes this path.What I found
The stored terminal matches the ingestion-failure write in
RunExecutionService, not a Cursor adapter terminal. ThecatchCausearound the provider-event consumer logsorchestration V2 provider event ingestion failedwith the real cause, then callswriteFinalRunEvents.makeProviderFailureonly keeps a specific message for a handful of error tags, so anything else (a schema error, for example) becomes the genericProvider turn failed.with classunknown, and the thread never shows the actual exception.writeFinalRunEventsthen runscascadeTerminalizeRunOwnedSubagents, which marks every still-open run-owned subagent, node and turn item failed in the same write. That explains the parent and four unfinished subagents failing together, while the one that had just completed was already settled.The Cursor run keeps going because that path never cancels it:
- Cursor sessions don't implement
subscribeEvents, so closingeventSubscriptiondoes nothing. interruptTurnis what callsrun.cancel, and the failure path doesn't call it.activeTurnstays set until the SDK run actually finishes.- The continued protocol frames mean the SDK callback is still succeeding; those events just queue up with nobody reading them.
This also likely explains the failed retry. The synthetic terminal makes the thread look reusable, but the next send hits
CursorAdapterV2.startTurn's guard (Cursor provider turn … is still active) before anything cancels the old run. Stop isn't offered either, because the run is already markedfailed. For now, quitting the desktop app closes the Cursor session, which is the practical way to drop the leftover turn.#15316 is the same ingestion-failure finalization gap with a Claude
AskUserQuestiontrigger. #12374, #13284 and #9047 are different paths.Likely fix area
RunExecutionService's ingestion-failure path. Some options:- Interrupt or cancel the provider turn when ingestion fails before the turn ends (for Cursor,
run.cancel). - Keep the original cause somewhere that survives log rotation, or surface it on the failed terminal, so the triggering event can be identified. Cancelling alone won't tell us which event threw.
A maintainer will decide on the fix direction.
- Cursor sessions don't implement
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 4, 2026
Before submitting
Area
apps/server
Steps to reproduce
Observed sequence, rather than a deterministic minimal reproduction:
The initial exception is not available in the retained logs, so I cannot identify the exact input/event that triggers this yet.
Expected behavior
T3's run status should agree with the provider's lifecycle. If event ingestion fails and T3 terminates the app run, it should cancel the underlying provider run and its owned work, or recover the event subscription. The original exception should remain available for diagnosis.
Actual behavior
The thread displayed
Provider turn failed.and recorded a terminal failure, while the underlying Cursor run continued for at least another 71 minutes. T3 also marked four unfinished native subagent tasks failed at the same instant.This makes it difficult to know whether work has stopped and risks overlapping work if the user retries. A later retry also failed to start, but its underlying cause was not recovered, so I am not claiming it was definitively caused by the lingering run.
Impact
Major degradation or frequent failure
Version or commit
0.0.46-nightly.20261003.2638Environment
7.2.6-201.nobara.fc44.x86_64@cursor/sdkfor this run path2026.08.04-aaa8809(version/help/auth status work; this is not evidence that the SDK uses that binary)Logs or stack traces
Sanitized timeline from the local orchestration database and provider logs; all times UTC:
There were 12,896 provider protocol updates for that same native run after the app failure, counted across the affected provider log and its two rotated files. No
run.completedfor that native run was found in those retained files. The app failure occurred at 19:07 PDT on October 3.No OS OOM kill or associated core dump was found in the inspected records. This is an app/provider lifecycle mismatch; it does not establish whether there was transient memory pressure before the failure.
Suspected failure path (not a confirmed original exception)
Inspection of the installed server bundle shows the provider-event consumer's
catchCausepath inRunExecutionService:orchestration V2 provider event ingestion failed.writeFinalRunEventswith a failed terminal of classunknownandopenRunOwnedSubagents.eventSubscription.That path matches the persisted failure shape and the continued native traffic. However, the warning with the original cause was not found in the retained logs, so the exact triggering exception remains unconfirmed.
Related reports
AskUserQuestionschema error. This report adds evidence from Cursor with native subagents; the Claude-specific trigger has not been established here.RetriableError: [canceled] http/2 stream closed with error code CANCEL (0x8)#12374 is a different Cursor transport failure on an earlier version.Workaround
No reliable workaround verified. A system restart did not prevent a subsequent provider failure. No settings or source changes were made during this investigation.
Raw logs, prompts, credentials, and project contents are intentionally omitted. This report was prepared with Codex from local read-only diagnostics and posted at the affected user's request.