Repository navigation
Give the Deep Agents harness the browse_cli and stagehand_code surfaces - #3158
sydney-runkle wants to merge 6 commits into
Conversation
`prepareBrowseCliHarnessAdapter` is already shared by claude_code and codex, but both of them reach the pinned wrapper through a shell. Two small openings let a harness without a shell reuse it unchanged: - expose `wrapperPath` so a caller can exec the wrapper directly instead of relying on it being first on PATH; - split the eval-harness addendum's invocation paragraph out of the skill builder, and export `buildBrowseSkillDocument` so a harness with no Skill tool can inline the same (single-source, non-drifting) skill text. No behavior change for claude_code or codex.
browse_cli has only ever been reachable from a shell, which limits it to the CLI-agent harnesses. This bridge serves the same surface over stdio MCP: one `browse` tool taking the exact command line the browse skill documents, run against the harness's pinned wrapper. Transport mirrors the Stagehand facade bridge — the MCP server lives in the harness process (so it logs through EvalLogger and shares the adapter's wrapper/env) and the agent spawns a dependency-free `node -e` relay to a loopback port. Safety is unchanged: commands go through the same `isAllowedBrowseCommand` allowlist claude_code enforces in `canUseTool`, and are tokenized rather than handed to a shell, so there is no interpreter to escape from.
DEEPAGENTS_TOOL_SURFACES listed only the MCP/facade surfaces, so every leaderboard number for deepagents comes from surfaces that cost one model round-trip per atomic browser action. claude_code is additionally benchmarked on browse_cli and the code-mode surfaces, which batch many browser steps per model call — a tool-surface asymmetry, not a harness-quality one. browse_cli is the half of that gap the Deep Agents runner can take today: it owns its own daemon rather than a CoreTool agent mount, so it short-circuits before `startAgentToolRuntime` exactly like the claude_code and codex adapters, and the pinned wrapper is fronted by the browse_cli MCP bridge (stdio MCP being the runner's only tool channel). The browse skill is inlined into the prompt instructions because there is no Skill tool to load it. The code-mode surfaces stay unsupported: their mounts arrive `via: "handles"` as live in-process JS objects, which a separate Python process cannot bind without an RPC bridge of their own.
`executeCodeExposureSnippet` lived privately in claudeCodeToolAdapter, with a copy in piToolAdapter carrying the note "consolidate when a third harness needs it". deepagents is the third, so it moves to framework/codeExposure.ts with the log category as a parameter — the only thing that differed between the copies. Scope semantics (handle names plus startUrl/task/console, bound by name) are unchanged, and both existing callers keep their own log categories.
The browse_cli bridge's transport — a harness-hosted MCP server behind a loopback socket, reached by a dependency-free `node -e` relay — is not specific to browse_cli. Splitting it out lets a second mount reuse it; what each bridge exposes stays with the bridge. No behavior change: the relay script, server lifecycle, and the spec handed to the agent are the same.
stagehand_code mounts `via: "handles"` — live in-process objects (the Stagehand client, its page) — so the snippet must execute in the harness process while the agent that wrote it runs elsewhere. That is the same split codex makes: codex fronts the executor with a loopback HTTP bridge plus a workspace client script its shell invokes, and deepagents, having no shell, fronts the same executor with the harness run tool on the loopback MCP bridge. Structural parity with the codex adapter: - the handles branch sits inside the same try as the MCP branch, with `bridge` declared alongside `cwd` so the setup-failure path tears both down; - cleanup closes the bridge, then the runtime, then removes the cwd; - an ObservationRecorder probes after every run — success or failure — so each run step consumes exactly one observation index. Two Deep Agents specifics: langchain-mcp-adapters reports tool names without a server prefix, so the agent sees a bare `run` and surfaces' references to `mcp__stagehand_browser__run` are rewritten the way mastra already rewrites them; and `recordObservation` is deliberately omitted on this path, since the bridge already probes on execution and recording from the runner's tool_result stream too would double-count every step. playwright_code and cdp_code mount identically and would work through the same path, but are left out until they have validation runs of their own.
|
|
This PR is from an external contributor and must be approved by a stagehand team member with write access before CI can run. |
The asymmetry
packages/evalsrestricted the Deep Agents harness tostagehand_facade,stagehand_facade_legacy,playwright_mcpandchrome_devtools_mcp— every oneof which costs a model round-trip per atomic browser action.
claude_codeandcodexadditionally getbrowse_cliand the code-mode surfaces, which batch manybrowser steps into one model call. Deep Agents was the only agentic harness
limited to one round-trip per action, which likely accounts for part of the
leaderboard gap (~74% / 567s vs ~89% / 326s).
This adds
browse_cliandstagehand_codeto that harness.Why it needed more than a longer array
The Deep Agents runner is a separate Python process whose only tool channel is
stdio MCP, and the eval profile excludes its
execute/filesystem builtins — sothere is no shell to run a
browsewrapper from, and no in-process scope to hosta handles mount.
browse_clialso produces no agent mount at all, which is whyclaude_code and codex branch before
startAgentToolRuntime.Both surfaces are therefore delivered over MCP:
mcpLoopbackBridge.ts— harness-hosted MCP server plus a dependency-free relay,extracted so both mounts share it.
browseCliMcpBridge.ts— onebrowsetool, validated by the sameisAllowedBrowseCommandallowlist claude_code uses, tokenized andexecFiled,no shell.
deepagentsCodeBridge.ts— aruntool executing snippets via the sharedexecutor, modelled on
codexCodeBridge(same out-of-process problem), probingafter every run including failures.
codeExposure.ts— the snippet executor, consolidated;claude_codeandpinow import it instead of keeping their own copies.
Notes for review
claude_code/pichanges are that consolidation, not behaviour changes.recordObservationstays MCP-only, matching the other harnesses: claude_code andcodex record no observations for
browse_clieither.runandthe surface's
mcp__stagehand_browser__runreferences are rewritten — sameapproach as mastra.
playwright_code/cdp_codeare the identical path and one line each, but areleft out as unvalidated.
Verification
Both bridges driven through the actual Python MCP client the runner uses, without
model credentials:
stagehand_coderan a snippet against a live Stagehand browserand returned a screenshot observation;
browse_clireturnedbrowse/0.11.1andrefused
rm -rf /. Typecheck,build:cli,build:esm,oxfmtandoxlintareclean. Tests: 1042 passed, 1 failed —
mastraRunner, which fails identically on aclean tree.
🤖 Generated with Claude Code