Skip to content

Give the Deep Agents harness the browse_cli and stagehand_code surfaces - #3158

Open
sydney-runkle wants to merge 6 commits into
browserbase:mainfrom
sydney-runkle:deepagents-browse-cli-tool-surface
Open

sydney-runkle wants to merge 6 commits into
browserbase:mainfrom
sydney-runkle:deepagents-browse-cli-tool-surface

Conversation

@sydney-runkle

@sydney-runkle sydney-runkle commented Oct 9, 2026 •

Copy link
Copy Markdown

The asymmetry

packages/evals restricted the Deep Agents harness to stagehand_facade,
stagehand_facade_legacy, playwright_mcp and chrome_devtools_mcp — every one
of which costs a model round-trip per atomic browser action. claude_code and
codex additionally get browse_cli and the code-mode surfaces, which batch many
browser steps into one model call. Deep Agents was the only agentic harness
limited to one round-trip per action, which likely accounts for part of the
leaderboard gap (~74% / 567s vs ~89% / 326s).

This adds browse_cli and stagehand_code to that harness.

Why it needed more than a longer array

The Deep Agents runner is a separate Python process whose only tool channel is
stdio MCP, and the eval profile excludes its execute/filesystem builtins — so
there is no shell to run a browse wrapper from, and no in-process scope to host
a handles mount. browse_cli also produces no agent mount at all, which is why
claude_code and codex branch before startAgentToolRuntime.

Both surfaces are therefore delivered over MCP:

  • mcpLoopbackBridge.ts — harness-hosted MCP server plus a dependency-free relay,
    extracted so both mounts share it.
  • browseCliMcpBridge.ts — one browse tool, validated by the same
    isAllowedBrowseCommand allowlist claude_code uses, tokenized and execFiled,
    no shell.
  • deepagentsCodeBridge.ts — a run tool executing snippets via the shared
    executor, modelled on codexCodeBridge (same out-of-process problem), probing
    after every run including failures.
  • codeExposure.ts — the snippet executor, consolidated; claude_code and pi
    now import it instead of keeping their own copies.

Notes for review

  • The claude_code/pi changes are that consolidation, not behaviour changes.
  • recordObservation stays MCP-only, matching the other harnesses: claude_code and
    codex record no observations for browse_cli either.
  • langchain-mcp-adapters drops server prefixes, so the agent sees a bare run and
    the surface's mcp__stagehand_browser__run references are rewritten — same
    approach as mastra.
  • playwright_code/cdp_code are the identical path and one line each, but are
    left out as unvalidated.

Verification

Both bridges driven through the actual Python MCP client the runner uses, without
model credentials: stagehand_code ran a snippet against a live Stagehand browser
and returned a screenshot observation; browse_cli returned browse/0.11.1 and
refused rm -rf /. Typecheck, build:cli, build:esm, oxfmt and oxlint are
clean. Tests: 1042 passed, 1 failed — mastraRunner, which fails identically on a
clean tree.

🤖 Generated with Claude Code

Stagehand Evals added 6 commits October 9, 2026 10:30
`prepareBrowseCliHarnessAdapter` is already shared by claude_code and codex,
but both of them reach the pinned wrapper through a shell. Two small openings
let a harness without a shell reuse it unchanged:

- expose `wrapperPath` so a caller can exec the wrapper directly instead of
  relying on it being first on PATH;
- split the eval-harness addendum's invocation paragraph out of the skill
  builder, and export `buildBrowseSkillDocument` so a harness with no Skill
  tool can inline the same (single-source, non-drifting) skill text.

No behavior change for claude_code or codex.
browse_cli has only ever been reachable from a shell, which limits it to the
CLI-agent harnesses. This bridge serves the same surface over stdio MCP: one
`browse` tool taking the exact command line the browse skill documents, run
against the harness's pinned wrapper.

Transport mirrors the Stagehand facade bridge — the MCP server lives in the
harness process (so it logs through EvalLogger and shares the adapter's
wrapper/env) and the agent spawns a dependency-free `node -e` relay to a
loopback port.

Safety is unchanged: commands go through the same `isAllowedBrowseCommand`
allowlist claude_code enforces in `canUseTool`, and are tokenized rather than
handed to a shell, so there is no interpreter to escape from.
DEEPAGENTS_TOOL_SURFACES listed only the MCP/facade surfaces, so every
leaderboard number for deepagents comes from surfaces that cost one model
round-trip per atomic browser action. claude_code is additionally benchmarked
on browse_cli and the code-mode surfaces, which batch many browser steps per
model call — a tool-surface asymmetry, not a harness-quality one.

browse_cli is the half of that gap the Deep Agents runner can take today: it
owns its own daemon rather than a CoreTool agent mount, so it short-circuits
before `startAgentToolRuntime` exactly like the claude_code and codex
adapters, and the pinned wrapper is fronted by the browse_cli MCP bridge
(stdio MCP being the runner's only tool channel). The browse skill is inlined
into the prompt instructions because there is no Skill tool to load it.

The code-mode surfaces stay unsupported: their mounts arrive `via: "handles"`
as live in-process JS objects, which a separate Python process cannot bind
without an RPC bridge of their own.
`executeCodeExposureSnippet` lived privately in claudeCodeToolAdapter, with a
copy in piToolAdapter carrying the note "consolidate when a third harness needs
it". deepagents is the third, so it moves to framework/codeExposure.ts with the
log category as a parameter — the only thing that differed between the copies.

Scope semantics (handle names plus startUrl/task/console, bound by name) are
unchanged, and both existing callers keep their own log categories.
The browse_cli bridge's transport — a harness-hosted MCP server behind a
loopback socket, reached by a dependency-free `node -e` relay — is not specific
to browse_cli. Splitting it out lets a second mount reuse it; what each bridge
exposes stays with the bridge.

No behavior change: the relay script, server lifecycle, and the spec handed to
the agent are the same.
stagehand_code mounts `via: "handles"` — live in-process objects (the Stagehand
client, its page) — so the snippet must execute in the harness process while
the agent that wrote it runs elsewhere. That is the same split codex makes:
codex fronts the executor with a loopback HTTP bridge plus a workspace client
script its shell invokes, and deepagents, having no shell, fronts the same
executor with the harness run tool on the loopback MCP bridge.

Structural parity with the codex adapter:
- the handles branch sits inside the same try as the MCP branch, with `bridge`
  declared alongside `cwd` so the setup-failure path tears both down;
- cleanup closes the bridge, then the runtime, then removes the cwd;
- an ObservationRecorder probes after every run — success or failure — so each
  run step consumes exactly one observation index.

Two Deep Agents specifics: langchain-mcp-adapters reports tool names without a
server prefix, so the agent sees a bare `run` and surfaces' references to
`mcp__stagehand_browser__run` are rewritten the way mastra already rewrites
them; and `recordObservation` is deliberately omitted on this path, since the
bridge already probes on execution and recording from the runner's tool_result
stream too would double-count every step.

playwright_code and cdp_code mount identically and would work through the same
path, but are left out until they have validation runs of their own.
@changeset-bot

changeset-bot Bot commented Oct 9, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: de91cd0

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

This PR is from an external contributor and must be approved by a stagehand team member with write access before CI can run.
Approving the latest commit mirrors it into an internal PR owned by the approver.
If new commits are pushed later, the internal PR stays open but is marked stale until someone approves the latest external commit and refreshes it.

@github-actions github-actions Bot added external-contributor Tracks PRs mirrored from external contributor forks. external-contributor:awaiting-approval Waiting for a stagehand team member to approve the latest external commit. labels Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

external-contributor:awaiting-approval Waiting for a stagehand team member to approve the latest external commit. external-contributor Tracks PRs mirrored from external contributor forks.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant