Skip to content

Browser drag timeout crashes the server and cancels unrelated threads #17322

Description

@AlexBabescu

What happened

Several Codex threads were working concurrently, then stopped together. The user also saw part of the conversation log disappear and felt that the session had reverted.

The remote T3 Code server crashed at 2026-10-08 21:16:21 UTC after a browser automation drag timed out. The background service restarted, and startup recovery cancelled four running threads. Recent messages and activity entries remain in the database. Partial history loss has not been confirmed.

Diagnosis

The service log ends with an uncaught-looking Playwright locator.dragTo TimeoutError, immediately followed by Active child exited unexpectedly (1). The systemd journal confirms exit status 1 at 21:16:21 UTC. Startup recovery at 21:16:29 reports terminalizedRuns: 4.

The likely error path is in drag in ServerBrowserPage.ts, lines 374 onward. It starts source.dragTo, then awaits the target's bounding box and a pointer update before awaiting the drag promise. A drag timeout during those intervening awaits can reject before a rejection handler is attached. This is a source-based explanation consistent with the crash, not a confirmed isolated reproduction.

The same drag implementation remains in the current main branch at the time of triage. No shipped fix was identified.

Steps to reproduce

Observed in a live install. A deterministic reproduction has not been run.

  1. Run several Codex threads on a Linux T3 Code background service.
  2. Have a thread perform a browser preview drag whose target cannot be resolved before the timeout.
  3. In this incident, the source was canvas[aria-label='draw image preview'] and the target was .advanced-editor > .advanced-heading strong, with a 5000 ms timeout.
  4. The service exited with status 1, restarted, and recovery cancelled four running threads.

Expected behavior is for the drag tool to return an error while the server and unrelated threads keep running.

Version

Crashing background service: 0.0.46-nightly.20261008.2819, commit 5e2225671f705fcd33f1ea5591b79ba612fb6974.

Desktop client: 0.0.46-nightly.20261008.2833.

Environment

Remote server: Linux x86_64, kernel 7.0.0-38-generic, systemd user service. Codex provider. Desktop client on macOS arm64, Darwin 27.0.0, connected over SSH. The local triage CLI was 0.0.45 under Node v26.8.2; those are not the crashing server's version or a measurement of its embedded Node runtime.

Evidence

Times are UTC. Home paths are redacted.

locator.dragTo: Timeout 5000ms exceeded.
  - waiting for locator('canvas[aria-label=\'draw image preview\']')
  - waiting for locator('.advanced-editor > .advanced-heading strong')
    at drag (<T3_HOME>/runtime/versions/0.0.46-nightly.20261008.2819/t3:320933:26)
    at async <T3_HOME>/runtime/versions/0.0.46-nightly.20261008.2819/t3:322582:13
  name: 'TimeoutError'
[service-launcher] Active child exited unexpectedly (1).

21:16:21 systemd: t3code.service: Main process exited, code=exited, status=1/FAILURE
21:16:26 systemd: Started t3code.service - T3 Code server.
21:16:29.386 V2 orchestration recovery completed
  terminalizedRuns: 4
  stoppedSessions: 1
  closedRequests: 0
  retiredEffects: 0
  requeuedEffects: 0

Read-only database inspection confirms four runs changed to cancelled at 21:16:29. Their threads retain 10 to 16 messages and 187 to 522 activity entries each. No rollback or deletion event for these four threads was found around the crash.

Related issues

Fix applied or workaround

With the user's approval, stopped and disabled the extra older background service using systemctl --user disable --now t3code.service. The desktop-managed .2833 server on port 3774 remained running and returned HTTP 200. This removes the competing-server condition, but does not fix the drag crash. No source patch or direct database write was made.

Filed by

Codex, GPT-6, via t3 triage.

Activity

  1. juliusmarminge commented on Oct 8, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the write-up. I read this against main at 9b0df1358. I did not reproduce the drag timeout.

    The drag promise can reject before anything observes it

    drag starts source.dragTo(...) and only awaits that promise after the target bounding box and a pointer update (ServerBrowserPage.ts:374-388):

    const dropped = source.dragTo(target, { timeout });
    const end = await target.boundingBox({ timeout }).catch(() => null);
    if (end) await pointer({ x: end.x + end.width / 2, y: end.y + end.height / 2 }, "move");
    await dropped;

    dragTo and boundingBox run together. If the drag times out while boundingBox or the pointer update is still pending, that rejection has no handler yet. Node treats a rejection as unhandled when nothing is attached by the end of that turn, so a later await dropped does not take it back. The same function on the crashing build (5e2225671f) has this shape, and it is unchanged on current main.

    I did not find another Playwright call in this file that is started and then left alone across an await. click also starts its click before doing other work, but Promise.race attaches to that promise in the same turn (ServerBrowserPage.ts:243-253). The caller of drag does await it (ServerBrowser.ts:1675-1678), which does not cover the gap inside drag.

    How that becomes a process exit

    There is no general handler that keeps the server up. Loading the server imports the Cursor driver, which imports cursorSdk.ts, and that module installs an unhandledRejection listener as a side effect (server.ts:57 → ProviderInstanceRegistryHydration.ts:37 → builtInDrivers.ts:27 → CursorDriver.ts:47 → CursorAuth.ts:19 → cursorSdk.ts:50-65). That listener ignores one Cursor shell-spawn failure. For every other rejection, when it is the only listener, it prints the reason and calls process.exit(1). A Codex session still loads this module, because the import is on the server boot path rather than on the Codex driver. The engines field asks for Node 22 or newer (package.json), where an unhandled rejection also exits if no listener is installed. Either way the child exits 1.

    The launcher then throws Active child exited unexpectedly (1) (serviceLauncher.ts:602) and the CLI prints that as [service-launcher] ... before setting exit code 1 (cli/serviceLauncher.ts:20-24). The name: 'TimeoutError' line in the report looks like console.error of the Playwright error, which is what this listener does. I have not seen the journal itself.

    The four cancelled threads look like startup recovery

    After the process comes back, startup runs provider-runtime recovery and logs V2 orchestration recovery completed (serverRuntimeStartup.ts:518-539). Recovery collects runs in preparing, starting, running, or waiting and writes each as cancelled (ProviderRuntimeRecoveryService.ts:71-80, :319-327). Sessions that are not already stopped or in error are set to stopped. The events I read update that runtime state. They do not delete thread messages. terminalizedRuns: 4 in the report matches that count, so the four cancellations look like a consequence of the exit, not a second bug in the drag tool.

    Related, not duplicates

    • #16794 is the same class of failure (one unhandled rejection exits the process and recovery stops other sessions) with a different exception, a provider pipe ECONNRESET.
    • #14115 is two servers on one database. That does not explain the drag exception.
    • #16921 is a browser timeout that evicts the shared browser host. It does not go through this un-awaited dragTo.

    No open PR changes this gap. #16941, #16728, and #16963 edit ServerBrowserPage.ts, and each of those heads still starts dragTo and then awaits the bounding box before awaiting the drag.

    Smallest fix

    Observe the drag promise in the same turn it is created, then rethrow after the pointer update. For example, attach .catch immediately, keep the pointer update, and throw the stored error afterwards. A new process-wide unhandledRejection listener is a worse place to fix this: the Cursor guard skips its process.exit(1) when another listener is already registered (cursorSdk.ts:56-58), so an extra listener can change exit behavior for unrelated rejections.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions