Repository navigation
Add API to list / execute tools? #16
Description
Activity
@bwalderman's points about session management in https://github.com/webmachinelearning/webmcp/pull/19/files intersects with the idea here.
If it's only the in-browser Agent interacting with the site at a time then it's easier for the site to reason about tools which are async. For instance, fetching some data and updating the UI incrementally. But if multiple Agents could interact with the site (like the browser's Agent and also an Agent powered by an extension) then the 2 could stomp on each other. Very likely a rare edge case but worth thinking about?
If it's only the in-browser Agent interacting with the site at a time then it's easier for the site to reason about tools which are async. For instance, fetching some data and updating the UI incrementally. But if multiple Agents could interact with the site (like the browser's Agent and also an Agent powered by an extension) then the 2 could stomp on each other. Very likely a rare edge case but worth thinking about?
Thinking about this, and #20, it might be helpful to have some kind of lock (similar to Pointer Lock) that only 1 user or agent can hold at a time. The user can interrupt any agent at any time by clicking a "Take control" button in their browser to retake the lock. An agent can only take the lock by requesting it from the user, and only if another agent isn't already running. Elicitation would let an agent yield the lock back to the user temporarily.
We also need an API / listener that agent can subscribe to, to get notifications about updated list of tools. In MCP this is:
{ "jsonrpc": "2.0", "method": "notifications/tools/list_changed" }The available tools change based on context of the page. For example, when you're filling out a form, the inputs change dynamically based on provided information, and we don't want WebMCP integrations sitting in a loop polling for changes.
We also need an API / listener that agent can subscribe to, to get notifications about updated list of tools.
The API design issue mentioned
registerToolandunregisterTool- that could directly communicate tools changed - but I think following MCP spec makes the most sense.Reacted by Ilya GrigorikWe haven't really explored what the integration for a non-browser Agent with WebMCP looks like. The Agent could be interacting with web pages via extension APIs or chrome devtools protocol (for automation). There's a couple of options:
-
The Agent executes script on the web page and discovers tools using the same Web APIs through which the site is declaring them to the browser.
-
The API surface the Agent is using (an extension API or chrome devtools protocol) provides higher level hooks to connect WebMCP with the Agent.
This issue was filed assuming approach 1) but I don't think that will work long term. We'll likely be adding a lot to the engine for Web specific adaption of MCP that the Agent shouldn't need to implement. For example, there will likely be a need for an attribute where the embedder can indicate that only read-only tools should be executed for content in an iframe.
It's better for the engine to apply these policies, the approach in 2).
-
RESOLUTION: The group looks into higher-level hooks to connect WebMCP with external agents for listing tools. This reduces coupling with MCP and subsequently browser implementation complexity. Javascript injection by external Agents to interact with WebMCP is not supported.
Closing this since there's no other use-case for this API. Filed #32 to follow up on the resolution above.
Reacted by Anssi Kostiainen
We didn't include it in our initial explainer since these are things the author ca already do, however, having an API to list out the registered tools and be able to execute them by name and argument dictionary would be useful for external agents (e.g. provided via extensions or third-party libraries).
Perhaps something like:
Would return an object like:
Which a third-party agent could use via
await window.agent.tools[0].execute({text: "foo"})