Skip to content

Tool calls longer than 60 s close a healthy launched browser and lose every page #2933

Description

@coygeek

Description of the bug

On main, every tool call is capped at 60 seconds, and when the cap fires the server treats the browser as dead: it closes the launched Chrome, starts a new one, and every open page with its console and network history is lost (with --isolated, the profile's cookies and storage go with the closed browser; that part was inferred, not inspected). The cap fires for valid calls on a healthy browser whose own parameters ask for longer work: wait_for with timeout: 70000, navigate_page with timeout: 65000, and an evaluate_script that awaits a 65-second promise. The failing call's error says the browser connection "may have been lost", which is not what happened. The released 1.10.1 package keeps the browser and every page for the same calls: evaluate_script returns its result, and wait_for fails only after its own 70-second timeout. The cap came from "fix: fail fast on tool calls stuck after a dead browser connection (#2699)" and is listed in the pending 1.11.0 release PR #2823, so the regression would ship with that release.

Actual behavior

On the main build each long call returns isError: true after 60.0 s with the same text, and the next call reports a browser restart with only a fresh blank page. Case A:

wait_for            isError=true  60004 ms
Tool "wait_for" timed out after 60000ms waiting on the browser connection. The connection may have been lost (for example, the debugged browser or app restarted). It will be re-established automatically on the next tool call.
list_pages          isError=false  1169 ms
Note: the browser was restarted or reconnected since the last call. Page ids have changed. Call list_pages to see open pages.
## Pages
3: about:blank [selected]

Case B (navigate_page) failed after 60005 ms and Case C (evaluate_script) after 60006 ms with the same message, each followed by the same restart note and a page list of 3: about:blank [selected]. Pages 1 and 2 were gone in every case. In other main runs of this investigation that hit the same cap (a stalled take_snapshot and a stalled evaluate_script), the client also recorded the Chrome browser process ID changing during the call, for example from 62321 to 67725.

On the release build the same inputs leave the browser alone. Case A returned isError=true after 70006 ms with Error: Timed out after waiting 70000ms, and Case C returned "done after 65s" after 65218 ms. In both, the next list_pages still showed 1: about:blank and 2: Real Title (http://127.0.0.1:PORT/titled.html) [selected]. In Case B on the release build, navigate_page had not returned when the client stopped waiting at 120 s; the list_pages request queued behind it was answered about 135 s after the navigation started and still listed page 2 (2: http://127.0.0.1:PORT/titled.html [selected]) with no restart note. How far navigate_page overshoots its timeout is a separate problem; the point here is that the release build kept the browser and its pages.

The same handler path serves the chrome-devtools CLI daemon. A separate CLI run from an earlier run in this investigation (navigate_page with --timeout 120000 to a page that responds after 75 s) left the Chrome process that existed before the call dead afterwards, and the next list_pages printed the restart note.

Reproduction

  1. Start a loopback fixture server whose /hang path never responds:
// fixture.mjs — run with `node fixture.mjs`; prints the port
import http from 'node:http';
const s = http.createServer((req, res) => {
  if (req.url === '/hang') return; // never responds
  res.writeHead(200, {'content-type': 'text/html'});
  res.end('<!doctype html><title>Real Title</title><h1>Titled</h1>');
});
s.listen(0, '127.0.0.1', () => console.log(s.address().port));
  1. Start the server under test with the flags above, initialize an MCP session, and send these tools/call requests in order (PORT is the fixture port). Each case below used a fresh server process.

Case A, wait_for with a 70-second timeout:

{"name":"new_page","arguments":{"url":"http://127.0.0.1:PORT/titled.html"}}
{"name":"list_pages","arguments":{}}
{"name":"wait_for","arguments":{"pageId":2,"text":["Never appears"],"timeout":70000}}
{"name":"list_pages","arguments":{}}

Case B, navigate_page with a 65-second timeout to a slow server:

{"name":"new_page","arguments":{"url":"http://127.0.0.1:PORT/titled.html"}}
{"name":"navigate_page","arguments":{"pageId":2,"url":"http://127.0.0.1:PORT/hang","timeout":65000}}
{"name":"list_pages","arguments":{}}

Case C, evaluate_script awaiting a 65-second promise:

{"name":"new_page","arguments":{"url":"http://127.0.0.1:PORT/titled.html"}}
{"name":"evaluate_script","arguments":{"pageId":2,"function":"() => new Promise(r => setTimeout(() => r('done after 65s'), 65000))"}}
{"name":"list_pages","arguments":{}}

Each case was run once per build (six runs); the main build failed at 60.0 s in all three of its runs. The cap is a fixed timer, so the failure is expected on every run whose call outlasts it.

Expectation

A valid, slow call on a healthy browser either finishes or fails on its own terms, and the browser with its pages stays as it was. The timeout parameter of wait_for, navigate_page and new_page is documented as "Maximum wait time in milliseconds. If set to 0, the default timeout will be used." (docs/tool-reference.md; timeoutSchema in src/tools/ToolDefinition.ts:487), with no upper bound in the schema, and the evaluate_script description gives async () => await fetch("example.com") as an example of an awaited function. Nothing in the tool reference, README.md or docs/ mentions a 60-second limit on a tool call or says that exceeding one discards the browser. The cap's own comment (src/ToolHandler.ts:29-40 on the main build) says it exists for a CDP transport that "dies silently", not for slow work on a live connection. The release build keeps the browser and its pages in all three cases.

MCP configuration

--headless --isolated --no-usage-statistics --no-performance-crux unless the reproduction states otherwise (update checks disabled with CHROME_DEVTOOLS_MCP_NO_UPDATE_CHECKS=1).

Chrome DevTools MCP version

1.10.1 (npm) and main 5ddb0a3 (local build)

Chrome version

154.0.8037.98 (stable)

Node version

v22.23.2

Operating system

macOS 27.0

Environment details

  • chrome-devtools-mcp 1.10.1 from npm (the release build).
  • chrome-devtools-mcp built from main at commit 5ddb0a3 (the main build; its --version still prints 1.10.1).
  • Google Chrome 154.0.8037.98 stable, headless; Node v22.23.2; macOS 27.
  • Both builds were started with --headless --isolated --no-usage-statistics --no-performance-crux and CHROME_DEVTOOLS_MCP_NO_UPDATE_CHECKS=1, and driven by a scripted stdio MCP client that sends one tools/call at a time and records wall time.

Evidence

  • Expected source: docs/tool-reference.md timeout parameter text for wait_for, navigate_page and new_page, and timeoutSchema in src/tools/ToolDefinition.ts:487; evaluate_script description example; the scope comment above TOOL_CALL_TIMEOUT_MS in src/ToolHandler.ts:29-40.
  • Failure source: scripted MCP runs of Cases A, B and C on the main build, with the release build as the control; the cap is TOOL_CALL_TIMEOUT_MS = 60_000 (src/ToolHandler.ts:40), applied to the whole handler and response by #raceWithTimeout (src/ToolHandler.ts:212-235, called for the handler and response at src/ToolHandler.ts:283-331), whose timeout callback calls forgetBrowserOnTimeout, which reaches BrowserManager.forget() (src/BrowserManager.ts:440-454) and closes a launched browser.
  • Evidence provenance: observed
  • Local verification: reproduced
  • Reproduction completeness: complete

All three cases were observed on both builds for this report; the CLI variant is from an earlier run in this investigation and was not rerun.

Fix check

With the input of Case C, evaluate_script returns "done after 65s" and the following list_pages still lists page 2 without a restart note. With Case A, wait_for fails only after its own 70-second timeout with the wait-timeout error, and page 2 is still listed afterwards. As a control that keeps the intent of #2699, a call that is stuck because the CDP transport is actually gone still fails within a bounded time and the next call reconnects.

Additional context

The full open and closed issue and PR corpus was screened as of 2026-10-06 (including searches for timed out after 60000ms, TOOL_CALL_TIMEOUT and browser was restarted); no issue reports this. The #2699 PR discussion covers the dead-transport case only. #2778 and PR #2805 concern the auto-connect consent timeout, and #2675 is an --autoConnect list_pages timeout on Windows; neither involves a launched browser being discarded after a slow but healthy call.

Other tools whose legitimate work can exceed 60 s, such as lighthouse_audit on a slow page, performance_start_trace with reload and autoStop, or take_heapsnapshot on a large page, go through the same handler path; they were not run for this report.

Two other observed failures reach this path on the main build only because they stall for longer than 60 s: evaluate_script called while a dialog is already open, and take_snapshot on a page with a pending navigation or a crashed renderer. On the release build those stall for Puppeteer's 180-second protocol timeout instead, so they are separate defects that this issue's fix does not resolve.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions