Skip to content

[Bug]: Cursor run continues after T3 marks parent and native subagents failed #15447

Description

@j0nsh1n

Before submitting

  • I searched existing issues and did not find an exact duplicate.
  • I included enough detail to investigate the problem; a deterministic minimal reproduction is not yet available.

Area

apps/server

Steps to reproduce

Observed sequence, rather than a deterministic minimal reproduction:

  1. Use the Cursor provider in the T3 Code desktop nightly listed below.
  2. Start an agent task that launches native subagents and leave it running.
  3. In the affected run, T3 marked the parent and four unfinished native subagent tasks failed immediately after another native subagent completed.
  4. Inspect the persisted orchestration events and Cursor provider protocol logs: the same native Cursor run keeps emitting updates after T3 has marked it failed.

The initial exception is not available in the retained logs, so I cannot identify the exact input/event that triggers this yet.

Expected behavior

T3's run status should agree with the provider's lifecycle. If event ingestion fails and T3 terminates the app run, it should cancel the underlying provider run and its owned work, or recover the event subscription. The original exception should remain available for diagnosis.

Actual behavior

The thread displayed Provider turn failed. and recorded a terminal failure, while the underlying Cursor run continued for at least another 71 minutes. T3 also marked four unfinished native subagent tasks failed at the same instant.

This makes it difficult to know whether work has stopped and risks overlapping work if the user retries. A later retry also failed to start, but its underlying cause was not recovered, so I am not claiming it was definitively caused by the lingering run.

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261003.2638

Environment

  • T3 Code desktop app, Linux x86_64, Nobara/Fedora 44, KDE Wayland
  • Kernel 7.2.6-201.nobara.fc44.x86_64
  • Cursor provider; installed server bundle uses @cursor/sdk for this run path
  • Cursor CLI on PATH: 2026.08.04-aaa8809 (version/help/auth status work; this is not evidence that the SDK uses that binary)

Logs or stack traces

Sanitized timeline from the local orchestration database and provider logs; all times UTC:

2026-10-04T02:07:06.054Z  native subagent node.completed
2026-10-04T02:07:06.060Z  subagent.completed
2026-10-04T02:07:06.064Z  turn-item.completed
2026-10-04T02:07:06.423Z  app parent run + 4 unfinished native subagent tasks marked failed
                         message: "Provider turn failed."
                         class: "unknown", code: null, retryable: null
2026-10-04T03:18:11.141Z  last observed update for the SAME native Cursor run

There were 12,896 provider protocol updates for that same native run after the app failure, counted across the affected provider log and its two rotated files. No run.completed for that native run was found in those retained files. The app failure occurred at 19:07 PDT on October 3.

No OS OOM kill or associated core dump was found in the inspected records. This is an app/provider lifecycle mismatch; it does not establish whether there was transient memory pressure before the failure.

Suspected failure path (not a confirmed original exception)

Inspection of the installed server bundle shows the provider-event consumer's catchCause path in RunExecutionService:

  1. Logs orchestration V2 provider event ingestion failed.
  2. Calls writeFinalRunEvents with a failed terminal of class unknown and openRunOwnedSubagents.
  3. Closes eventSubscription.
  4. Does not call provider interrupt/cancel in that catch path.

That path matches the persisted failure shape and the continued native traffic. However, the warning with the original cause was not found in the retained logs, so the exact triggering exception remains unconfirmed.

Related reports

Workaround

No reliable workaround verified. A system restart did not prevent a subsequent provider failure. No settings or source changes were made during this investigation.

Raw logs, prompts, credentials, and project contents are intentionally omitted. This report was prepared with Codex from local read-only diagnostics and posted at the affected user's request.

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful timeline and bundle notes, @j0nsh1n. This is a real lifecycle bug, and it's still present on current main (737993303d). No commit since your nightly changes this path.

    What I found

    The stored terminal matches the ingestion-failure write in RunExecutionService, not a Cursor adapter terminal. The catchCause around the provider-event consumer logs orchestration V2 provider event ingestion failed with the real cause, then calls writeFinalRunEvents. makeProviderFailure only keeps a specific message for a handful of error tags, so anything else (a schema error, for example) becomes the generic Provider turn failed. with class unknown, and the thread never shows the actual exception.

    writeFinalRunEvents then runs cascadeTerminalizeRunOwnedSubagents, which marks every still-open run-owned subagent, node and turn item failed in the same write. That explains the parent and four unfinished subagents failing together, while the one that had just completed was already settled.

    The Cursor run keeps going because that path never cancels it:

    • Cursor sessions don't implement subscribeEvents, so closing eventSubscription does nothing.
    • interruptTurn is what calls run.cancel, and the failure path doesn't call it. activeTurn stays set until the SDK run actually finishes.
    • The continued protocol frames mean the SDK callback is still succeeding; those events just queue up with nobody reading them.

    This also likely explains the failed retry. The synthetic terminal makes the thread look reusable, but the next send hits CursorAdapterV2.startTurn's guard (Cursor provider turn … is still active) before anything cancels the old run. Stop isn't offered either, because the run is already marked failed. For now, quitting the desktop app closes the Cursor session, which is the practical way to drop the leftover turn.

    #15316 is the same ingestion-failure finalization gap with a Claude AskUserQuestion trigger. #12374, #13284 and #9047 are different paths.

    Likely fix area

    RunExecutionService's ingestion-failure path. Some options:

    • Interrupt or cancel the provider turn when ingestion fails before the turn ends (for Cursor, run.cancel).
    • Keep the original cause somewhere that survives log rotation, or surface it on the failed terminal, so the triggering event can be identified. Cancelling alone won't tell us which event threw.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions