Skip to content

Parallel function calls: when one tool raises, finished siblings lose their responses and may be re-run聽#7428

Description

@thuongvu

馃敶 Required Information

Describe the Bug:

When the model makes parallel function calls and one FunctionTool raises (no on_tool_error_callback), ADK cancels
the other calls and re-raises. That's fine. But a sibling that had already finished isn't saved to the session
either. On the next turn, drop_orphaned_function_calls strips it from the request, so the model never learns that
the call ran. If the tool had a side effect (created a ticket, sent a message), the model is likely to do it again.

I think the cause is that _execute_prepared_function_calls (flows/llm_flows/tools/_batch_executor.py) only builds
the merged response event after _gather_or_cancel returns, so the finished calls' responses are thrown away with the
batch when one call raises.

Steps to Reproduce:

  1. pip install google-adk==2.11.0
  2. Run the script below (no API key; the model is a scripted BaseLlm that calls both tools in parallel).
  3. create_ticket returns at once; notify_oncall raises 200 ms later. The script prints what the session kept.

Expected Behavior:

The session keeps create_ticket's real response, so the next turn shows the model that the ticket exists. The
failing call gets no response, and nothing is made up. ADK already does this in two similar places: the confirmation
pause (#6732, 67d3e49) and workflow nodes (2c759e9, "keep outputs of sibling nodes that finished in the same tick as
a failing node").

Observed Behavior:

raised: pager service returned 500
tickets created: ['disk full']
responses in session: []

The ticket was created, but the session has no response for it. On the next turn the call is dropped
(Dropping function calls with no matching function response: [...]), so the model has no record of it.

Environment Details:

  • ADK Library Version (pip show google-adk): 2.11.0; also main at fac77be
  • Desktop OS: macOS 15
  • Python Version (python -V): 3.12.13

Model Information:

  • Are you using LiteLLM: Not in the repro (scripted BaseLlm); also seen through LiteLlm, see below.
  • Which model is being used: N/A for the repro; gpt-5-mini in the end-to-end check below.

馃煛 Optional Information

Regression: Not checked.

Additional Context:

Minimal Reproduction Code:

import asyncio

from google.adk.agents import LlmAgent
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_response import LlmResponse
from google.adk.runners import InMemoryRunner
from google.genai import types

tickets = []  # stands in for an external system with side effects


async def create_ticket(title: str) -> dict:
  tickets.append(title)
  return {"ticket_id": "T-1"}


async def notify_oncall(message: str) -> dict:
  await asyncio.sleep(0.2)  # create_ticket has finished by now
  raise RuntimeError("pager service returned 500")


class ScriptedModel(BaseLlm):  # calls both tools in parallel
  async def generate_content_async(self, llm_request, stream=False):
    call = types.Part.from_function_call
    yield LlmResponse(content=types.Content(role="model", parts=[
        call(name="create_ticket", args={"title": "disk full"}),
        call(name="notify_oncall", args={"message": "disk full"}),
    ]))


async def main():
  agent = LlmAgent(name="ops", model=ScriptedModel(model="scripted"),
                   tools=[create_ticket, notify_oncall])
  runner = InMemoryRunner(agent=agent, app_name="repro")
  session = await runner.session_service.create_session(app_name="repro", user_id="u")
  message = types.Content(role="user", parts=[types.Part(text="Open a ticket and page on-call.")])
  try:
    async for _ in runner.run_async(user_id="u", session_id=session.id, new_message=message):
      pass
  except RuntimeError as e:
    print("raised:", e)
  session = await runner.session_service.get_session(app_name="repro", user_id="u", session_id=session.id)
  print("tickets created:", tickets)
  print("responses in session:", [r.response for e in session.events for r in e.get_function_responses()])


asyncio.run(main())

How often has this issue occurred?:

  • Always (100%): every run on 2.11.0.

Activity

  1. surajksharma07 commented on Oct 7, 2026

    @surajksharma07
    Collaborator

    @thuongvu Reproduced on 2.11.0 and also on 1.34.1: responses in session come back empty and if I send a follow-up turn the scripted model calls create_ticket a second time.

    Until #7429 is reviewed a workaround worth trying: set on_tool_error_callback on the agent to return something like {"error": str(error)}. The failing call then gets a normal response, the batch doesn't raise and on my side both responses end up in the session so there's no second ticket. Could you check whether that holds up with your LiteLlm/gpt-5-mini setup? It does change behaviour though since the error no longer reaches the caller so it may not suit you if you rely on the exception.

    For #7429 could you also run the tests against a resume after a node retry so we can be sure the kept results aren't saved twice? That will help when maintainers review it.

  2. thuongvu commented on Oct 8, 2026

    @thuongvu
    Author

    @surajksharma07 Thanks for reproducing.

    Could you check whether that holds up with your LiteLlm/gpt-5-mini setup?

    The workaround works. With on_tool_error_callback returning {"error": str(error)} with LiteLlm/gpt-5-mini, create_ticket's result was kept and the model didn't redo it. Yes, like you said, the error doesn't reach the caller.

    Could you also run the tests against a resume after a node retry so we can be sure the kept results aren't saved twice? That will help when maintainers review it.

    I tried it, and the finished create_ticket call ran once and its result was saved once. Added the test to this commit on the PR.

  3. surajksharma07 commented on Oct 8, 2026

    @surajksharma07
    Collaborator

    Thanks for confirming the workaround and adding the retry/resume test @thuongvu. Pulled fde7957 and the 16 targeted tests pass. Also applied the src changes on top of current main and the repro keeps {'ticket_id': 'T-1'} once with the original error still raised.

    One thing before review: the branch now conflicts with main in tests/unittests/test_runners.py (main changed it after fac77be, the source files still merge cleanly). Could you rebase it so it's ready when the workflows get approved?

    Will flag #7429 for maintainer review. Until it's merged and released the on_tool_error_callback workaround is the way to go.

  4. thuongvu commented on Oct 9, 2026

    @thuongvu
    Author

    @surajksharma07 Thank you! I rebased.

  5. surajksharma07 commented on Oct 9, 2026

    @surajksharma07
    Collaborator

    The rebase looks good @thuongvu. 3e309d5 now merges cleanly into current main and with it the targeted tests (30) plus the tools/workflow/runner suites pass on my end. The repro keeps {'ticket_id': 'T-1'} once and still raises the original error.

    Nothing else needed from you for now. It's ready for the maintainers once the 2 pending workflows get approved and flagged it for review.

    Until it's merged and released the on_tool_error_callback workaround is still the way to go. Will update here if review comes back with anything.

  6. stringsofthemind-oss commented on Oct 9, 2026

    @stringsofthemind-oss

    Sharing a separate, reproducible execution-safety test inspired by this issue.

    We've built a disposable Google ADK integration lab in Once:
    stringsofthemind-oss/once#304

    The lab compares released ADK 2.11.0, the exact base of #7429, and the proposed fix at 3e309d5. It uses a local HTTP provider with an independent SQLite effect journal and no native deduplication.

    Observed results:

    • On the affected release/base, history-based replay creates two raw provider effects; the proposed ADK fix reduces this to one.
    • When replay is deliberately forced, both the affected and proposed ADK versions create two raw effects.
    • With Once protecting the same logical operation identity and effect-bearing inputs, the tested replay scenarios retain one provider effect.
    • Fresh-process tests also exercise confirmed replay, changed-input CONFLICT, durable UNKNOWN after lost acknowledgement/crash, and authoritative fixture reconciliation.

    Preserving successful tool results in agent history and preventing repeated external effects are complementary safeguards. The proposed ADK fix addresses the first problem; durable operation identity can help address the second.

    The reproduction is pinned, credential-free and independently checkable. Windows/Linux CI, release tests, clean-checkout reproduction and negative controls have passed.

    Instructions and evidence:
    https://github.com/stringsofthemind-oss/once/blob/29a21ea7b6506b06b533cb267294ad3c765953f1/examples/adk-parallel-replay/README.md

    This is deterministic local-fixture evidence, not a live-model reproduction, production-provider qualification or Google-endorsed integration.

    We'd welcome an independent reproduction or adversarial counterexample, especially one that breaks the identity, replay, UNKNOWN or provider-effect-count assertions.

    No changes to Google ADK are being requested here; this is intended as a complementary safety test.

  7. added
    spam[Status] Issues suspected of having comments which are spam
    on Oct 10, 2026
  8. adk-bot commented on Oct 10, 2026

    @adk-bot
    Collaborator

    馃毃 Automated Spam Detection Alert 馃毃
    @maintainers, a suspected spam comment was detected in this thread.

    Reason:

    @stringsofthemind-oss posted promotional content promoting a 3rd party product/repository ("Once") with external links.
    
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

spam[Status] Issues suspected of having comments which are spamtools[Component] This issue is related to tools

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions