Skip to content

Add API to list / execute tools? #16

Description

@bokand

We didn't include it in our initial explainer since these are things the author ca already do, however, having an API to list out the registered tools and be able to execute them by name and argument dictionary would be useful for external agents (e.g. provided via extensions or third-party libraries).

Perhaps something like:

window.agent.tools[0];

Would return an object like:

{
    name: "add-todo",
    description: "Add a new todo item to the list",
    inputSchema: {
        type: "object",
        properties: {
            text: { type: "string", description: "The text of the todo item" }
        },
        required: ["text"]
    },
    execute: <...function object...>
}

Which a third-party agent could use via await window.agent.tools[0].execute({text: "foo"})

Activity

  1. khushalsagar commented on Sep 5, 2025

    @khushalsagar
    Collaborator

    @bwalderman's points about session management in https://github.com/webmachinelearning/webmcp/pull/19/files intersects with the idea here.

    If it's only the in-browser Agent interacting with the site at a time then it's easier for the site to reason about tools which are async. For instance, fetching some data and updating the UI incrementally. But if multiple Agents could interact with the site (like the browser's Agent and also an Agent powered by an extension) then the 2 could stomp on each other. Very likely a rare edge case but worth thinking about?

  2. bwalderman commented on Sep 9, 2025

    @bwalderman
    Collaborator

    If it's only the in-browser Agent interacting with the site at a time then it's easier for the site to reason about tools which are async. For instance, fetching some data and updating the UI incrementally. But if multiple Agents could interact with the site (like the browser's Agent and also an Agent powered by an extension) then the 2 could stomp on each other. Very likely a rare edge case but worth thinking about?

    Thinking about this, and #20, it might be helpful to have some kind of lock (similar to Pointer Lock) that only 1 user or agent can hold at a time. The user can interrupt any agent at any time by clicking a "Take control" button in their browser to retake the lock. An agent can only take the lock by requesting it from the user, and only if another agent isn't already running. Elicitation would let an agent yield the lock back to the user temporarily.

  3. igrigorik commented on Sep 20, 2025

    @igrigorik

    We also need an API / listener that agent can subscribe to, to get notifications about updated list of tools. In MCP this is:

    {
      "jsonrpc": "2.0",
      "method": "notifications/tools/list_changed"
    }
    

    The available tools change based on context of the page. For example, when you're filling out a form, the inputs change dynamically based on provided information, and we don't want WebMCP integrations sitting in a loop polling for changes.

  4. jasonjmcghee commented on Sep 21, 2025

    @jasonjmcghee

    We also need an API / listener that agent can subscribe to, to get notifications about updated list of tools.

    The API design issue mentioned registerTool and unregisterTool - that could directly communicate tools changed - but I think following MCP spec makes the most sense.

  5. khushalsagar commented on Oct 1, 2025

    @khushalsagar
    Collaborator

    We haven't really explored what the integration for a non-browser Agent with WebMCP looks like. The Agent could be interacting with web pages via extension APIs or chrome devtools protocol (for automation). There's a couple of options:

    1. The Agent executes script on the web page and discovers tools using the same Web APIs through which the site is declaring them to the browser.

    2. The API surface the Agent is using (an extension API or chrome devtools protocol) provides higher level hooks to connect WebMCP with the Agent.

    This issue was filed assuming approach 1) but I don't think that will work long term. We'll likely be adding a lot to the engine for Web specific adaption of MCP that the Agent shouldn't need to implement. For example, there will likely be a need for an attribute where the embedder can indicate that only read-only tools should be executed for content in an iframe.

    It's better for the engine to apply these policies, the approach in 2).

  6. anssiko commented on Oct 2, 2025

    @anssiko
    Member

    RESOLUTION: The group looks into higher-level hooks to connect WebMCP with external agents for listing tools. This reduces coupling with MCP and subsequently browser implementation complexity. Javascript injection by external Agents to interact with WebMCP is not supported.

  7. khushalsagar commented on Oct 2, 2025

    @khushalsagar
    Collaborator

    Closing this since there's no other use-case for this API. Filed #32 to follow up on the resolution above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions