Skip to content

[Bug]: After a preview automation timeout, the session fails over to a client on another machine that doesn't have the thread open #13540

Description

@QarthO

Filed by Claude (Opus 5.5, running inside T3 Code) on behalf of @QarthO, at his request. The investigation and reproduction were done by the agent on his machine; he confirmed what he saw on screen at each step.

Before submitting

Area

apps/server (preview automation broker), with visible effects in apps/web

Summary

When a preview automation request times out, the broker disconnects the host and the agent's session fails over to another connected client. That other client can be a desktop app on a different computer that doesn't have the thread open. It then loads the preview there, so local-only URLs (localhost, 127.0.0.1, *.orb.local, etc.) fail with "This site can't be reached". Because tab state is shared, the failure also shows up in the preview the user is actually watching. From the user's side, the preview looks like it breaks every time the agent touches it.

Steps to reproduce

  1. Run T3 Code on machine A (the server runs here too).
  2. Open the desktop app on machine B and connect it to machine A's server. Leave it on the new-thread screen; don't open the thread.
  3. On machine A, open a thread and a preview tab on a URL that only resolves on machine A (for example http://localhost:3000, or an OrbStack https://myapp.orb.local).
  4. Cause a preview automation timeout on machine A. The easiest way is [Bug]: preview_resize always times out (and evicts the automation host) when the main window is zoomed, because the guest webview reports innerWidth × hostZoom #12319: zoom the UI in once (View > Zoom In), then have the agent call preview_resize with a freeform size. It times out after 15 s.
  5. Have the agent call preview_evaluate or preview_navigate on the same tab.

Expected behavior

After the timeout, the agent keeps controlling the preview on machine A (the one showing the thread), or gets a clear error. It shouldn't silently move to a client on another machine that doesn't have the thread open.

Actual behavior

  • preview_evaluate returns chrome-error://chromewebdata/, and its results don't match what machine A shows. The page it reports has a different creation time from the one on screen.
  • preview_navigate reports success, but the page is an error page. On machine A the preview flashes the real page for a split second, then shows "This site can't be reached … ERR_NAME_NOT_RESOLVED" (for 127.0.0.1 URLs it's the same, pointing at machine B's own localhost).
  • Clicking Reload on machine A always works immediately, but the agent's next preview action breaks it again.
  • Public URLs (e.g. https://example.com) work, which makes it look like a DNS or network problem on machine A. It isn't: curl and other browsers on machine A reach the site fine.
  • It persists until machine B's app is closed. After closing it, the same tab immediately responded with machine A's page, and the same preview_resize (with the zoom reset) succeeded straight away.

Evidence it's machine B handling the commands: machine A's desktop.trace.ndjson has no PreviewManager.automationEvaluate / navigate spans for any of the agent's calls during the failure. As soon as machine B's app was closed, they appear again. The server trace shows the handover at the timeout:

21:59:12  preview_resize (freeform 992x800)      -> times out (machine A's UI zoomed)
21:59:27  PreviewAutomationBroker.closeConnection / disconnect
21:59:28  PreviewAutomationBroker.acquireConnection / connect   (machine A reconnects)
          ... every later preview call goes to machine B until its app is closed

Why it happens (from reading apps/server/src/mcp/PreviewAutomationBroker.ts)

  1. On timeout, awaitResponse disconnects the host (as described in [Bug]: Optional 500 ms preview metadata timeout disconnects the automation host #12273 / [Bug]: preview_resize always times out (and evicts the automation host) when the main window is zoomed, because the guest webview reports innerWidth × hostZoom #12319).
  2. The next request looks for a host in the environment. Hosts that own the target tab are sorted first, but a host without the thread or tab is still eligible. If machine A hasn't reconnected yet (it came back about a second later here), machine B is the only candidate and gets the assignment.
  3. The assignment is then pinned: "a live assignment … is not silently moved to a newer client". So even after machine A reconnects, is focused and owns the visible tab, the session stays on machine B. (fix(preview): use the visible browser for new agent sessions #13064's "prefer the client showing the tab" only applies when there's no live assignment, and at the moment of failover machine A wasn't connected to be preferred.)
  4. Machine B's preview host loads the tab URL on its own machine and network. The resulting load failure is written to the shared tab state, which is why machine A's panel shows the error too.

Possible solutions

  • Only fail over to hosts that have the thread's preview open. Treat ownsTargetTab as a requirement rather than a sort key. If no host qualifies, return a clear "no host has this preview open" error instead of picking any client.
  • Let a reconnecting host reclaim its session (the follow-up fix(preview): use the visible browser for new agent sessions #13064 left out). If the host that just lost its assignment reconnects within a short grace period, or if a focused host owns the visible target tab while the pinned host doesn't, move the assignment back instead of keeping it pinned to the other client.
  • Make the handover visible. Include which client or device handled a command in tool results and errors (today it's only failed on client preview-<id>). Also consider keeping one client's load failure out of the shared tab state that other clients render.

Fixing #12319 and not evicting hosts on timeouts (#12273) would remove the usual trigger, but any other timeout can still cause this handover.

Impact

Major degradation or frequent failure. The agent can't use the preview at all, and it looks like the agent is breaking the user's preview, which is confusing to debug.

Version or commit

T3 Code (Nightly) 0.0.43-nightly.20260924.2223 (Electron 44.4.2, Chrome 152) on machine A. Machine B was the Windows desktop app.

Environment

Machine A: macOS, Apple Silicon, running the server and the desktop app; the dev site is served by OrbStack at an *.orb.local HTTPS address. Machine B: Windows desktop app connected to machine A's server over the network, idle on the new-thread screen. Agent provider: Claude Code.

Activity

  1. juliusmarminge commented on Sep 25, 2026

    @juliusmarminge
    Member

    Triage: #13540

    Verdict: Confirmed bug, still present on main. Not a duplicate. No newer commit fixes it.

    Severity: High for any environment with a second desktop connected. After one automation timeout, later preview tools run on the other machine and the shared tab is overwritten with that machine's load failure. The agent then acts on a different page from the one on screen.

    User version: Nightly 0.0.43-nightly.20260924.2223, which already includes #13064 (merged 2026-09-24). PreviewAutomationBroker.ts has not changed since that PR. Current main (d23eab13d9) has the same selection logic.

    What is wrong

    A preview-automation timeout disconnects the host and deletes that session's lease. The next call is treated as unassigned and may be given to any other client in the environment, including a desktop that does not have the thread open. That choice is then pinned for the life of the other client's connection. The original machine reconnecting, focusing, and owning the visible tab does not take the session back.

    The other client loads the tab URL on its own network. localhost, 127.0.0.1, and *.orb.local fail there. It reports that failure into the shared preview session, so the machine the user is watching covers the real page with "This site can't be reached". Reload on the watching machine fixes the shared snapshot until the next agent preview call, which is still pinned to the other machine.

    Why the code does this

    On an unanswered request, the broker disconnects that host and completes its stream so it can reconnect. Disconnect drops every lease on that connection:

          const result = yield* Deferred.await(deferred).pipe(Effect.timeoutOption(timeoutMs));
          return yield* Option.match(result, {
            onNone: () =>
              Effect.gen(function* () {
                // An unanswered request invalidates this connection. Do not replay
                // actions: the client may have applied them before becoming unreachable.
                yield* disconnect(connection.clientId, connection.queue, true);
                return yield* new PreviewAutomationTimeoutError(requestContext);
              }),

    With no live lease, the next call sorts every host in the environment. Owning the tab is only a sort key. A host with no liveTabs is still eligible, and the winner is stored as the new lease:

          // Keep one provider session on one physical desktop runtime so a
          // multi-step browser interaction cannot jump between independent
          // Electron cookie/DOM state. A live assignment that predates an
          // operation is not silently moved to a newer client: the caller gets a
          // capability failure and can deliberately start a fresh provider
          // session. A dead lease is pruned above and may fail over.
          const ownsTargetTab = (host: ClientConnection, visibleOnly = false) =>
            host.liveTabs.some(
              (tab) =>
                tab.threadId === input.scope.threadId &&
                (!visibleOnly || tab.visible === true) &&
                (input.tabId === undefined || tab.tabId === input.tabId),
            );
          const connection =
            hasLiveAssignment && supportsOperation(assignedConnection, input.operation)
              ? assignedConnection
              : hasLiveAssignment
                ? undefined
                : Array.from(current.clients.values())
                    .filter(
                      (host) =>
                        host.environmentId === input.scope.environmentId &&
                        supportsOperation(host, input.operation),
                    )
                    .sort(
                      (left, right) =>
                        Number(ownsTargetTab(right, true)) - Number(ownsTargetTab(left, true)) ||
                        Number(ownsTargetTab(right)) - Number(ownsTargetTab(left)) ||
                        Number(right.focused) - Number(left.focused) ||
                        right.focusOrder - left.focusOrder,
                    )[0];

    A client on the new-thread screen reports liveTabs: [] until it is asked to run a command. Its window still counts as focused when it is the foreground window (document.hasFocus() and visibilityState === "visible"). Two windows can both be focused, one per machine.

    That leaves two ways the idle machine wins:

    1. The next call arrives while the timed-out host is still reconnecting. The idle client is the only candidate and is pinned.
    2. The original host has reconnected but has not reported focus yet. Registration sets focused: false and liveTabs: []. If the other window is focused, it sorts first. The pin then sticks after the original host reports the visible tab.

    #13064 only changes selection when there is no live lease, and it says it does not move a session already pinned to another browser. This is that remaining case. The tests lock it in: fails over a pinned provider session only after its host disconnects and evicts an unanswered host and lets later calls use a healthy runtime both expect the next call to go to a second host that never reported the tab.

    Why the watching preview shows the error

    The idle client does not refuse a tab it has never opened. It lists the shared session, waits for a local webview, and navigates that webview (PreviewAutomationHosts.tsx, navigate around the requireReadyTab path). liveTabs only includes tabs whose desktop state has web contents, so this client does not look like an owner until after it has already been pinned and has built its own webview.

    That webview's LoadFailed status is always reported. PreviewManager.reportStatus writes it onto the shared session with no check of which client owns the visible tab. Every client applies the failed event. PreviewView then hides the real surface and shows PreviewUnreachable. Reload goes through the local webview, which can still reach the site, and reports success. The next agent command is still on the other machine, so it fails again.

    preview_evaluate returning chrome-error://chromewebdata/ with a different document creation time matches a newly created webview on the other machine, not the page already open on the first. Public URLs working is the same path: the other machine can resolve those names.

    Related issues

    Issue Relationship
    #13051 / #13064 Fixed new-session routing to the visible tab. Explicitly does not move an existing pin. This report is that gap.
    #12319 Open. Zoomed preview_resize is a reliable way to hit the 15s broker timeout. Trigger only.
    #12273 Open. A 500ms follow-up status after other preview tools can also evict the host. Trigger only.
    #12146 Closed. Host failed to reconnect after eviction. Reconnect now works; this bug is where the session goes during that gap.
    #11167 Closed. Routing to a suspended mobile client. Different host class.

    Fixing the timeout triggers removes the usual way in. Any other unanswered preview request can still hand the session to the wrong machine.

    Fix direction

    On failover, do not select a host that does not own the thread's tab. If none does, return PreviewAutomationNoAvailableHostError and do not write a lease. The reconnecting owner can take the next call once focusHost reports liveTabs.

    Keep today's behavior for a new session that is opening the first tab: no host owns a tab yet, so the focused client should still be eligible. The dangerous case is an existing tab whose owner disappeared for the reconnect gap. input.tabId is usually set by the preview tools; a dropped lease currently forgets the last tab id, so retain it across the disconnect or treat "tab id present and nobody owns it" as no host.

    Do not move a live lease just because another client is focused. That was an intentional #13064 limit so cookies and DOM stay on one runtime. A narrow reclaim is still worth it for the already-pinned case: the pinned host does not have the tab visible, and another host does.

    Also stop applying one client's LoadFailed to the shared snapshot that every other client renders. Routing is the main fix; this is what makes the idle machine paint over the preview the user is looking at.

    Replace the two failover tests above. They currently require the bug. Add coverage for: second client connected with empty liveTabs, timeout, original host reconnects and reports the visible tab, later calls stay on the original host (or fail with no host during the gap), and the idle client never receives the command.

    Workaround

    Close the other desktop app. The lease is dropped with that connection, and the next call can select the machine that has the thread open. Quitting only the watching app does not help while the other one stays connected.

    Confidence

    High on the mechanism. It matches the reported trace shape (timeout, disconnect, reconnect about a second later, later commands absent from that machine's desktop trace until the other app exits) and the shared LoadFailed overlay. Not re-run on two desktops in this triage.

  2. added
    via-triageFiled through npx t3 triage
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    on Sep 25, 2026
  3. coygeek commented on Oct 3, 2026

    @coygeek

    Summary

    This corroborates #13540 rather than proposing a separate issue. With two macOS desktop clients connected to one T3 Code environment, a preview timeout was followed by later tool calls running on the other machine. The calls still supplied the same explicit tabId. A loopback URL that worked on the machine running the application then failed from the newly selected host.

    Steps to reproduce

    The following describes the recorded failure sequence with a generic site and anonymized hosts. It has not been repeated in a fresh two-host test during this investigation. The smallest missing reproduction detail is the host reconnect and focus timing needed to make the second host win after the timeout.

    1. Run T3 Code and a local test site on Host A. For example, serve a static page at http://localhost:5173 on Host A only.
    2. Connect a second desktop client on Host B to the same T3 Code environment.
    3. On Host A, open the test page in the thread's preview. Retain the returned tab ID and confirm the page loads.
    4. In the same agent session, call preview_wait_for for text that never appears with timeoutMs: 10000. In the recorded session, a preview_resize timeout also preceded the host change.
    5. After the timeout, call preview_status and preview_navigate for the same explicit tab ID. Keep using http://localhost:5173 as the generic navigation target.
    6. Compare the automation host handling these requests before and after the timeout, and confirm from Host A that its HTTP server is still reachable.

    Expected behavior

    Subsequent commands for an existing tab should keep its physical browser host, or return a clear temporary host-unavailable error while that host reconnects. A condition miss must not silently make the session control an independent browser on another machine. Retaining an explicit tab ID should retain the intended browser context.

    Actual behavior

    After the timeout, the automation client changed from Host A to Host B. Later calls with the same explicit tab ID used Host B's native browser. Navigation to the loopback service failed there while the original application remained reachable on Host A.

    One returned status had a saved viewport setting of 1280 x 800 and a measured viewport of 1843 x 1152. This mismatch was another signal that the tool state and the browser being controlled no longer matched. It does not establish the cause of the timeout or host change.

    Affected area

    apps/server/src/mcp/PreviewAutomationBroker.ts, with visible effects in the desktop preview and apps/web automation host. The failure is host affinity for an existing tab after a timeout. The wait and resize timeouts are observed triggers, not separate defects in this report.

    Runtime or environment

    Two macOS desktop clients connected to one environment, using T3 Code's collaborative preview tools. The exact desktop and server builds used for the recorded failure have not been established in this investigation. The upstream source below was inspected at main commit fed41fa88bb27cb4325cb208d571393850bc63c2; that is a source inspection revision, not a claim about the original running build.

    Evidence

    The recorded browser verification notes state that an ordinary page-condition timeout disconnected the assigned automation host and a subsequent call selected the other machine's native browser. The tool observations also distinguished the automation clients before and after the timeout. Original logs and screenshots are omitted because they contain unrelated private application context. Host labels and the test URL above are anonymized examples, not literal log values.

    The public broker source still has the mechanism identified in #13540. Its timeout path disconnects an unanswered host. Connection removal deletes its assignments. Host selection treats ownership of the target tab as a sorting preference, while another capable host in the environment remains eligible. A new live assignment is then retained across later operations. This verifies the source mechanism and matches the observations, but does not prove the exact reconnect timing of the recorded incident.

    Impact

    Major degradation during local browser verification. The agent can inspect or navigate a different physical browser from the one that can reach the local application, despite continuing to target the same tab ID.

    Additional context

    #13540 is the existing report for this failure and #13772 is an open proposed fix. #12898 covers a related preview_wait_for deadline race. This report should be added as corroborating evidence on #13540, not opened as a duplicate issue.

    A reachable temporary tunnel URL allowed verification to continue across the two hosts. That avoided the loopback reachability symptom; it did not verify that host affinity was restored.

    A regression check should use two hosts and an existing tab owned by Host A. After a condition timeout or an unanswered request, commands for that tab should either stay on Host A or fail clearly until it reconnects. Host B must not receive commands for the retained tab merely because Host A is reconnecting. Once Host A reports the tab again, the same session should resume there.

  4. janpilu commented on Oct 6, 2026

    @janpilu

    Posted by Claude (Opus 5.5, running inside T3 Code) on behalf of @janpilu, at their request. The reproduction was run by an agent on their machine.

    Still reproduces on Nightly 0.0.46-nightly.20261005.2667 (commit 37de6cbde65c), macOS 26.6.2 arm64, Electron 44.4.2. Two things this adds to the report:

    1. A trigger that needs no zoom or resize. Any request that hits the broker deadline does it. preview_evaluate with a promise that never resolves times out after 15 s every time.
    2. Recordings get split across hosts. preview_recording_stop goes to the other host, which has no recording for that tab. The recording itself is fine and is still running on the original host.

    Repro

    Two desktop clients were connected to one environment: A preview-f2725c8c… and B preview-ea4278c9…. The page was a throwaway local HTTP server on 127.0.0.1, with no auth.

    preview_open({ open: true, reuseExistingTab: false, show: false })          // -> tab_d
    preview_navigate({ tabId: "tab_d", url: "http://127.0.0.1:61606", readiness: "load" })
    // identity probe: an evaluate that throws names client A in its error
    // on A: window.t3ReproMarker = "client-A", document.cookie = "t3_repro=client-A"
    
    preview_recording_start({ tabId: "tab_d" })                                  // recording: true on A
    preview_evaluate({ tabId: "tab_d", expression: "new Promise(() => {})", awaitPromise: true })
    // -> "Preview automation evaluate timed out after 15000ms."
    preview_recording_stop({ tabId: "tab_d" })
    // -> "Preview automation recordingStop failed on client preview-ea4278c9…"   (B)

    After that, the same explicit tabId: "tab_d" reaches B:

    { "url": "chrome-error://chromewebdata/", "marker": null, "pageInstance": null }

    Server trace (server.trace.ndjson, UTC 2026-10-06)

    Time Span Note
    07:39:32.820 PreviewAutomationBroker.invoke evaluate, 15 001.7 ms, PreviewAutomationTimeoutError
    07:39:47.820 PreviewAutomationBroker.disconnect same trace as the timeout
    07:39:47.835 PreviewAutomationBroker.invoke recordingStop sent to B 15 ms after the disconnect
    07:39:48.827 PreviewAutomationBroker.acquireConnection A reconnects about 1 s later, too late

    B's underlying error, which the tool result reduces to "failed on client" (#15336):

    Preview automation request preview-1798 found no active recording for tab tab_d on environment <env> thread <thread>.
    

    To get the video back, I forced the same timeout on B. The next preview_recording_stop went back to A and returned the original recording: 3.26 MB, 56.5 s, 566 decoded frames. So the recording wasn't lost. The stop request just went to the wrong host.

    The lease-retention approach in #13772 looks like it covers this, because recordingStop would then only go to the host that owns tab_d. It may be worth adding a broker test for recording start → timeout → stop. We hit this in four separate verification threads on Oct 3 and Oct 5, and it's the most common reason our agents give up on the integrated browser.

    Separately, after the timeout, ordinary preview_evaluate calls on the affected host also kept timing out until a navigation cleared the stuck evaluate. That's consistent with #16264.

  5. WhiteWarrior625 commented on Oct 8, 2026

    @WhiteWarrior625

    Additional recorded recurrence on Linux, stock nightly 0.0.46-nightly.20261005.2702, with a local desktop and a desktop on another machine connected to the same environment. This is evidence for the pre-server-browser architecture, not a reproduction against current main.

    The same logical tab existed on both automation clients. Two 30-second navigate timeouts were followed by routing to the other client. An optional post-action status lookup then reached its 500 ms deadline and changed the selected host again. The remote desktop could not load the server machine's localhost URL. Agent-side tab metadata continued to change while the user's visible panel retained an earlier connection-refused page. The owner reported two more routing recurrences on October 8.

    Installed source confirms the shared mechanism: the provider-session lease names a client connection; timeout removes that client and lease; a later call can choose another connected desktop. The optional status lookup uses the same broker timeout path. The transport had also stalled, but we have not established the cause of that stall or claimed that client selection alone fixes it.

    I checked the newer stock 0.0.46-nightly.20261008.2819 source. It includes PR15328's environment-owned preferred browser and avoids host eviction for optional status reads. The downloaded stock artifact is retained with a matching GitHub SHA256; it has not been installed or verified live here. We will use a coordinated safe update boundary, rather than interrupt session-owned work.

    The remaining visibility distinction is also tracked by #17015. A successful agent navigation, shared URL/viewport metadata or visible: true cannot by itself certify the exact panel the user is watching. We require a check of the user-facing page and viewport before claiming a live walkthrough. No local runtime fork was installed.

    Nightly check on October 8: v0.0.46-nightly.20261008.2819, commit 5e2225671f705fcd33f1ea5591b79ba612fb6974, is still the newest published Nightly. Its packaged server confirms the environment-owned preferred browser and the non-evicting optional status path. The ordinary request timeout branch still disconnects the selected connection when updateCurrentTab is not false. That source path alone does not prove the old cross-machine failure recurs with the new browser architecture, so I am keeping the partial fix separate from live acceptance.

    The server browser status derives visibility from connected streamed viewers. That cannot identify the particular physical panel the user is watching. We still need a user-facing page/viewport check for walkthrough acceptance. Packaged server SHA256: 310635e55d11b09ef92672c2ea71d44f017cab43b26a83586b16be6ab051cfa5. No upgrade or local runtime patch was applied.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions