Skip to content

[Bug]: Provider turn interrupt failures are silently swallowed — Stop generation looks successful while the turn keeps running #4524

Description

@henrik092

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server (with a smaller follow-on in apps/web)

Summary

When the provider call behind Stop generation fails, the failure is only written to a log line. The orchestration command is still receipted as accepted, no thread activity is appended, and the UI shows nothing. From the user's perspective the Stop button is simply dead while the turn keeps running.

The codebase already has the right mechanism for this — it is just not wired to this branch. In processTurnInterruptRequested, the "no active provider session" case calls appendProviderFailureActivity({ kind: "provider.turn.interrupt.failed", ... }). The case where the provider request itself fails does not.

Steps to reproduce

  1. Start a long-running turn (long tool call, browser automation, etc.).
  2. Put the provider in a state where turn/interrupt is rejected, so providerService.interruptTurn fails with a ProviderAdapterRequestError. Any provider-side rejection reaches the same path — see "How I hit it" below for the concrete trigger I observed.
  3. Click Stop generation (repeatedly, if you like).
  4. Watch the thread UI, the activity feed, and ~/.t3/userdata/logs/server-child.log.

Expected behavior

A failed interrupt is surfaced to the user — a provider.turn.interrupt.failed activity in the thread, or a thread error — consistent with how the "no active provider session" branch already behaves.

Actual behavior

  • The command is receipted accepted with no error.
  • thread.turn-interrupt-requested is persisted normally.
  • The provider call fails.
  • The only trace is a logWarning; the thread gets no error activity and the UI stays silent.
  • The turn continues running.

Code path (upstream/main @ 41a430a)

  • apps/server/src/orchestration/Layers/ProviderCommandReactor.ts:907 — the interrupt is issued by session only (interruptTurn({ threadId })), so the runtime's internal turn id decides the target.
  • apps/server/src/orchestration/Layers/ProviderCommandReactor.ts:1072-1081 — processDomainEventSafely catches the cause and terminates it in Effect.logWarning("provider command reactor failed to process event", …). Nothing is emitted to the thread.
  • apps/web/src/components/ChatView.tsx:4830 — onInterrupt only calls setThreadError when the dispatch fails. A dispatch that succeeds and a provider call that later fails is indistinguishable from success in the UI.

Contrast with ProviderCommandReactor.ts:896-904, which does the right thing for the missing-session case.

Impact

Major degradation or frequent failure — when it happens the user has no working way to stop a running agent and no indication why.

Version or commit

upstream/main @ 41a430a88 (code paths above read directly from that ref). Observed at runtime on a downstream fork that is based on it.

Environment

Linux Mint 22.3, kernel 7.0.0-28-generic, Node 22.19.0, desktop AppImage build, Codex provider (gpt-5.6-sol), Codex app-server.

Logs or stack traces

# ~/.t3/userdata/logs/server-child.log — repeated 16x in 17s, one per click
{
  eventType: 'thread.turn-interrupt-requested',
  cause: 'ProviderAdapterRequestError: Provider adapter request failed (codex) for turn/interrupt: expected active turn id 019f98f9-… but found 019f98f7-…
      at mapCodexRuntimeError (…/apps/server/dist/bin.mjs:66909:9)'
}

Meanwhile, in state.sqlite:

-- 16 events, 16 distinct command ids, every receipt accepted, zero receipt errors
SELECT r.status, count(*), sum(r.error IS NOT NULL)
  FROM orchestration_events e
  JOIN orchestration_command_receipts r ON r.command_id = e.command_id
 WHERE e.event_type = 'thread.turn-interrupt-requested';
-- accepted|16|0

No provider.turn.interrupt.failed activity was appended to the thread in that window.

How I hit it, and why it may matter for #231

Full disclosure on the trigger: the specific mismatch above came from a turn/steer implementation in my own fork, not from upstream — upstream/main has no turn/steer call. So the root cause of my mismatch is mine to fix, and I am not asking you to fix that part.

The reason I am filing here anyway is that the swallow-the-failure path is entirely upstream code and will hide any provider interrupt rejection, whatever the cause.

Since #231 (feat: add Steer and Queue follow-up modes) is in progress, one concrete note that may save some time. In my fork the sequence was:

  1. turn/steer is called on the running turn with expectedTurnId: <current>.
  2. The response carries a different turn id than the turn the provider is actually still executing.
  3. That returned id is written straight into the runtime's activeTurnId.
  4. Every later interrupt reads that value and is rejected with expected active turn id X but found Y.

I confirmed step 2/3 independently: the id T3 sent decodes as a UUIDv7 minted at 11:11:31.027Z, 72 ms after the steer message — while the provider's own event stream stayed on the original turn id for the entire turn (2078 occurrences of the original id, 0 of the new one).

If the upstream steer implementation also adopts the steer response's turn id as the active one, it can land in the same state. Validating that response — or at least making the interrupt failure visible — would cover it.

Workaround

Send a plain text instruction ("stop working now, take no further actions"). That is cooperative, not a real interrupt, so an in-flight tool call can still finish first. In my incident the turn only ended that way.

Activity

  1. papatistos commented on Aug 23, 2026

    @papatistos

    Additional reproduction on T3 Code Alpha 0.0.33, macOS Apple Silicon, OpenCode 1.18.21.

    This reproduces the same failure with two concurrent sessions using the OpenCode adapter and model opencode/x-preview-f-free.

    Observed sequence:

    1. Both sessions returned repeated Provider finish_reason: network_error events between 2026-08-23 14:24:48 and 14:25:10 UTC.
    2. Clicking Stop created thread.turn-interrupt-requested events, but T3 then persisted both sessions as running again.
    3. The provider runtimes later reported stopped, while the T3 UI continued to show Working after a full app restart.
    4. The local session records showed activeTurnId still set and a running session despite the provider runtime being stopped. I had to repair the two local records manually so the turns became interrupted and the sessions became stopped.
    5. After restarting, pressing Continue on one session at 2026-08-23 15:04:08 UTC caused five more network-error retries and a final Runtime error at 15:06:04 UTC. The session became error, but the provider runtime still reported running.

    Expected behavior is that a failed provider abort or terminal provider error always clears activeTurnId, settles the turn, and transitions the session to stopped, ready, or error. The UI should surface the underlying provider error and should not require manual local-state repair.

    The original provider network failure may be an upstream trigger, but the stuck Working state and the inconsistent persisted runtime state are reproducible T3 behavior.

  2. t3dotgg commented on Aug 28, 2026

    @t3dotgg
    Member

    Note

    🤖 GPT-5.6 Sol responding on behalf of Theo

    Thanks for the detailed report. We believe the original rejected-interrupt path is fixed by PR #7412.

    The fix catches an interrupt failure, attempts provider-session teardown, clears the active turn, stores the error, and appends visible failure activity. The later OpenCode sequence is not proven to be the same rejected provider call. It still shows a live Stop and reconciliation defect after repeated network_error events, including stale running state after restart and the reverse state where T3 reports an error while the provider runtime reports running. Those remaining conditions are now preserved in open issue #4713.

    I'm closing this as partly fixed and consolidated as part of an automated pass on all open issues. The remaining OpenCode Stop and runtime-reconciliation bugs are not fixed, and issue #4713 remains open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions