Repository navigation
Elicitation #21
Description
Activity
From @khushalsagar's comment in explainers-by-googlers/script-tools#2, elicitation can happen in the tool call itself:
For example, the site could expose a makePayment tool which when invoked brings up the payment flow on the site that requires user interaction.
This is an ideal way to elicit information from the user in human-in-the-loop scenarios; just get them to do it manually the way they always have.
What's missing is a way for the agent to know that the elicitation is happening on-page so it can do things like display a "waiting" message in the chat, extend its timeout, etc.
What's missing is a way for the agent to know that the elicitation is happening on-page so it can do things like display a "waiting" message in the chat, extend its timeout, etc.
Agreed. I initially thought this is as simple as adding a new annotation to the tool:
needsUserInput, so the Agent can know whether a tool will require user input. But there will be cases where a tool can conditionally require user input. For example, thesearchProductstool needs the delivery zip code but only if the user hasn't set a default address yet. So we need some way for the site to indicate to the Agent when it's waiting on user input. Both for the Agent to reflect this in its UI and ensure the web page is in a state where user interaction can proceed (if the tab was backgrounded for example).There's 2 issues we need to think through when designing for this:
- Potential for abuse. The site just wants to grab user attention as much as possible. Maybe something we can learn from how popups deal with it.
- Reusing MCP concepts. WebMCP is trying to align with MCP to the extent possible but elicitation is where the control flow is fundamentally different. With MCP, the service is remote. It has to rely on the MCP client/Agent to seek user input. That's why the protocol involves
elicitation/createmessage from server to client. By design the MCP/client/Agent is aware of user input yielding/progress because the client is the one arbitrating it.
With WebMCP, it's the MCP server (the site) which is managing user input directly. So I don't think we can try to build upon MCP's elicitation API design here. We have to think through something Web specific.
@MiguelsPizza wdyt?
I think it makes sense to align with the MCP spec here and defer elicitation call resolutions to the client. If the website wants to elicit user input in the webpage but also have the user input be inputted into said webpage, they can do so via a promise in the body of an async tool call.
For example:
const formTool = window.mcp.registerTool('example_form_tool', { description: 'Collects user information', parameters: { type: 'object', properties: { message: { type: 'string' } } } }, async (params) => { // maybe we add something like this? await window.mcp.focusTab(); return new Promise((resolve) => { // the website handles the elicitation internally showFormModal(params.message).then(userInput => { resolve(userInput); }); }); });
The downside of the above is I'm not sure how we would draw attention to the page in the case where the website needed user input. I agree with @khushalsagar here, providing a
focusTabapi like in the example gives too much power to any individual bad actor website. Also, in-page user input makes sense for a client like MCP-B which navigates to the tab before executing a tool, but not so much for the background executions @bwalderman was proposing.One of the benefits of deferring elicitations to the client is that you can still have an agent fill out the elicitation request since it's up to the client for how to handle the elicitation events. This would be much better for non-human in the loop remote browser automation tasks where forcing a human input in a tab would require DOM parsing/screenshots to fill out.
The larger question we might need to discuss is what variety of clients and use cases do we want to tailor the WebMCP APIs to? Just targeting human-in-the-loop provides us much more flexibility than also trying to support full autonomous browser automation, but I'd argue supporting both will be important for broad adoption.
The larger question we might need to discuss is what variety of clients and use cases do we want to tailor the WebMCP APIs to? Just targeting human-in-the-loop provides us much more flexibility than also trying to support full autonomous browser automation, but I'd argue supporting both will be important for broad adoption.
We need to have a path where tool can signal to autonomous workflow that user input or review is required. With my commerce hat on, this is often due to terms, disclosures, liability & compliance requirements, etc. The agent can build a cart and prefill the information, but then the tool needs a mechanism to indicate that hand-off is required.
There's been some design thinking since the issue was filed, reiterating the decision from #25 below:
- WebMCP is the API between site <-> user agent (browser).
- There is another API between browser <-> agent to connect WebMCP to the agent.
- When possible, WebMCP tries to mirror MCP concepts. Since most agents understand MCP, it makes 2) easier.
My original thinking was that elicitation with WebMCP only needs 1). The browser can manage it all internally but I'm realizing there's more to it.
- One usage pattern is for the Agent to run a remote browser instance in some VM. The user only interacts with the browser when needed with some remote access setup. The Agent needs to know when the site is blocked on user input.
- The point raised here about the Agent knowing of user progress. I haven't dug through the MCP spec to understand the use-case for it.
Also, we haven't concluded on what a standardized API for 2) looks like. The same question came up at #16 (comment). Sounds like we need to get more concrete about the API for 2) to guide 1)?
RESOLUTION: Tool execution should be able to start/stop yielding to the user throughout its lifecycle.
Rough idea of what yielding to the user might look like in the imperative API:
// Math tool. async function execute(input, agent) { // Operands. const { a, b } = input; // User needs to pick an operation manually. const operation = await agent.requestUserInteraction(async () => { return await new Promise((resolve, reject) => { const form = document.querySelector("#form"); form.addEventListener('submit', e => { e.preventDefault(); const operation = form.querySelector('select#operation').value; resolve(operation); }); }); }); let output; switch (operation) { case "add": output = a + b; break; case "subtract": output = a - b; break; // etc.. } return `Calculation complete. Result: ${output}`; }
Tool functions get a second parameter, an object representing the calling agent. The tool code can call
agent.requestUserInteraction. This tells the browser to enable user input on the relevant tab if input is currently disabled, flash the tab header if the tab isn't in the foreground, etc. The user has control of the page until the passed in promise or async function resolves, and then the agent takes back control. Any return value from the inner promise/function is passed through and returned fromrequestUserInteractionso that the tool code can use it.Tool functions get a second parameter, an object representing the calling agent
This is good idea and very similar to the api in AssistantUI
@Yonom and learnings/thoughts to share here?
The API sounds good to me. Was thinking if we should have an option where
requestUserInteractioncan fail but I'm not sure what the site is expected to do in such a case. If a site is abusing this API for user attention, the best alternative for the browser is likely to disallow the use of these tools for that site.For an in-browser scenario,
requestUserInteractionI can't really think of a reason for requestUserInteraction to fail. It should automatically yield control to the user. I suppose a headless/autonomous agent using WebMCP might want to reject a user interaction request to indicate to the page that there's no user available. That's out of scope for now but could be handled by lettingrequestUserInteractionthrow/reject. Whatever Extension/DevTools API is eventually built to expose WebMCP to external agents will need a way to notify clients of interaction requests and a way to resolve/reject them.About abusing this API for attention... it should only be available when there's an active tool call in progress, so the user should have already started their agent and asked it to do something on the page. For a tab/window that doesn't have an agent connected and so has never had a tool call, there'd be no way to call this API.
Reacted by nocluemichealAbout abusing this API for attention... it should only be available when there's an active tool call in progress, so the user should have already started their agent and asked it to do something on the page. For a tab/window that doesn't have an agent connected and so has never had a tool call, there'd be no way to call this API.
I was considering the case where the site intentionally keeps calling this during legitimate tool calls. Say the tool is
addProductand the site says, "look at this offer first", requiring the user to hit a button to dismiss. The user should at least have the option to say, "this site keeps asking me to provide input for unrelated stuff".Reacted by nocluemichealOne idea that might help with the “attention abuse during a legitimate call” scenario is to make the yield itself observable
- When a tool invokes
requestUserInteraction(...), the browser emits a minimal lifecycle signal (e.g., yield start/end) tagged with the currentexecutionIdand tool identity. This lets the agent reflect “waiting” and extend timeouts, and gives the browser enough context to apply its own policies. - On the browser side, vendors could choose to rate-limit or downgrade attention during a single execution, and expose a user affordance like “mute further/all prompts from this site/tool for this task”, without changing the site API surface.
- This could also open a path to "preapproved actions" for non-human in the loop remote browser automation tasks.
Reacted by nocluemicheal- When a tool invokes
Thanks for the feedback!
When a tool invokes requestUserInteraction(...), the browser emits a minimal lifecycle signal (e.g., yield start/end) tagged with the current executionId and tool identity.
I didn't follow which entity these events are targeting. The proposed API provides them to the author implicitly, "yield start" is when the callback passed to requestUserInteraction is invoked and "yield end" is when the promise returned by that callback resolves.
I also didn't follow what
executionIdis referring to here since the proposed API doesn't have that concept yet. The browser needs to limit to one tool invocation from a single Agent at a time. And the tool identity is also implicit since user interaction is limited to the duration of that one tool exection.Reacted by Stalgia GriggReacted by nocluemichealI didn't follow which entity these events are targeting. The proposed API provides them to the author implicitly, "yield start" is when the callback passed to
requestUserInteractionis invoked and "yield end" is when the promise returned by that callback resolves.Thanks for clarifying and apologies for my misunderstanding. I was splicing a few different thoughts and didn't yet fully understand some fundamentals here. The implicit state covers everything important.
There is still an interesting kernel wrt the browser->agent bridge and disambiguating types of "yield" unique to WebMCP but that is out of scope for this conversation.
Reacted by nocluemichealSorry for poor connection on call.
I just wanted to raise the question of whether we're planning initial permission akin to microphone access.
If this is the case, the user has already opted in to some degree, so the popup style mitigation seems more reasonable.
But also wanted to call out I felt @bwalderman was pushing for requesting access up front allowing requesting user interaction vs @khushalsagar pushing for popup style mitigation, so wasn't sure we were quite ready to declare resolution at that point (@anssiko)
Reacted by nocluemichealRESOLUTION: requestUserInteraction API implementation should give user an option to block abusive sites permanently but throw an error to developers so legitimate sites can implement fallback behaviour.
Reacted by nocluemichealI just wanted to raise the question of whether we're planning initial permission akin to microphone access.
I wasn't thinking about it like microphone. Up front permission makes sense in that case, you're sharing private user data with the site. But I don't see why the user would need to consent to WebMCP tools and/or elicitation in these tools. There needs to be some consent on whether a model can take agentic actions on a site but how those actions happen isn't a detail the user should worry about.
But because eliciation can be abusive, there will need to be some dialog for the user to say, "don't allow this site" or "always allow" and popups is the closest parallel that comes to mind.
That said, the API doesn't have to dictate how the user agent should manage these permissions. A microphone style opt-in is also fine, the API should be flexible such that the user agent can take any approach. And the resolution we landed on should enable both, do point out if it doesn't.
Reacted by nocluemicheal@khushalsagar i think the resolution makes sense
Reacted by nocluemichealClosing this since the resolved API has been added to the explainer.
Note to self: We should add IDLs for all the interfaces as well.
Reacted by nocluemicheal
Gathering thoughts on supporting MCP elicitation since this would be a good way to bring the user's attention to a tab if the agent determines that their input is needed.
This has also been discussed previously at explainers-by-googlers/script-tools#2