Repository navigation
Parallel function calls: when one tool raises, finished siblings lose their responses and may be re-run聽#7428
Description
Activity
@thuongvu Reproduced on 2.11.0 and also on 1.34.1: responses in session come back empty and if I send a follow-up turn the scripted model calls create_ticket a second time.
Until #7429 is reviewed a workaround worth trying: set on_tool_error_callback on the agent to return something like {"error": str(error)}. The failing call then gets a normal response, the batch doesn't raise and on my side both responses end up in the session so there's no second ticket. Could you check whether that holds up with your LiteLlm/gpt-5-mini setup? It does change behaviour though since the error no longer reaches the caller so it may not suit you if you rely on the exception.
For #7429 could you also run the tests against a resume after a node retry so we can be sure the kept results aren't saved twice? That will help when maintainers review it.
@surajksharma07 Thanks for reproducing.
Could you check whether that holds up with your LiteLlm/gpt-5-mini setup?
The workaround works. With
on_tool_error_callbackreturning{"error": str(error)}with LiteLlm/gpt-5-mini, create_ticket's result was kept and the model didn't redo it. Yes, like you said, the error doesn't reach the caller.Could you also run the tests against a resume after a node retry so we can be sure the kept results aren't saved twice? That will help when maintainers review it.
I tried it, and the finished
create_ticketcall ran once and its result was saved once. Added the test to this commit on the PR.Thanks for confirming the workaround and adding the retry/resume test @thuongvu. Pulled fde7957 and the 16 targeted tests pass. Also applied the src changes on top of current main and the repro keeps {'ticket_id': 'T-1'} once with the original error still raised.
One thing before review: the branch now conflicts with main in tests/unittests/test_runners.py (main changed it after fac77be, the source files still merge cleanly). Could you rebase it so it's ready when the workflows get approved?
Will flag #7429 for maintainer review. Until it's merged and released the on_tool_error_callback workaround is the way to go.
@surajksharma07 Thank you! I rebased.
The rebase looks good @thuongvu. 3e309d5 now merges cleanly into current main and with it the targeted tests (30) plus the tools/workflow/runner suites pass on my end. The repro keeps {'ticket_id': 'T-1'} once and still raises the original error.
Nothing else needed from you for now. It's ready for the maintainers once the 2 pending workflows get approved and flagged it for review.
Until it's merged and released the on_tool_error_callback workaround is still the way to go. Will update here if review comes back with anything.
- addedtools[Component] This issue is related to tools[Component] This issue is related to tools
on Oct 9, 2026 Sharing a separate, reproducible execution-safety test inspired by this issue.
We've built a disposable Google ADK integration lab in Once:
stringsofthemind-oss/once#304The lab compares released ADK 2.11.0, the exact base of #7429, and the proposed fix at 3e309d5. It uses a local HTTP provider with an independent SQLite effect journal and no native deduplication.
Observed results:
- On the affected release/base, history-based replay creates two raw provider effects; the proposed ADK fix reduces this to one.
- When replay is deliberately forced, both the affected and proposed ADK versions create two raw effects.
- With Once protecting the same logical operation identity and effect-bearing inputs, the tested replay scenarios retain one provider effect.
- Fresh-process tests also exercise confirmed replay, changed-input CONFLICT, durable UNKNOWN after lost acknowledgement/crash, and authoritative fixture reconciliation.
Preserving successful tool results in agent history and preventing repeated external effects are complementary safeguards. The proposed ADK fix addresses the first problem; durable operation identity can help address the second.
The reproduction is pinned, credential-free and independently checkable. Windows/Linux CI, release tests, clean-checkout reproduction and negative controls have passed.
Instructions and evidence:
https://github.com/stringsofthemind-oss/once/blob/29a21ea7b6506b06b533cb267294ad3c765953f1/examples/adk-parallel-replay/README.mdThis is deterministic local-fixture evidence, not a live-model reproduction, production-provider qualification or Google-endorsed integration.
We'd welcome an independent reproduction or adversarial counterexample, especially one that breaks the identity, replay, UNKNOWN or provider-effect-count assertions.
No changes to Google ADK are being requested here; this is intended as a complementary safety test.
- addedspam[Status] Issues suspected of having comments which are spam[Status] Issues suspected of having comments which are spam
on Oct 10, 2026 馃毃 Automated Spam Detection Alert 馃毃
@maintainers, a suspected spam comment was detected in this thread.Reason:
@stringsofthemind-oss posted promotional content promoting a 3rd party product/repository ("Once") with external links.
馃敶 Required Information
Describe the Bug:
When the model makes parallel function calls and one
FunctionToolraises (noon_tool_error_callback), ADK cancelsthe other calls and re-raises. That's fine. But a sibling that had already finished isn't saved to the session
either. On the next turn,
drop_orphaned_function_callsstrips it from the request, so the model never learns thatthe call ran. If the tool had a side effect (created a ticket, sent a message), the model is likely to do it again.
I think the cause is that
_execute_prepared_function_calls(flows/llm_flows/tools/_batch_executor.py) only buildsthe merged response event after
_gather_or_cancelreturns, so the finished calls' responses are thrown away with thebatch when one call raises.
Steps to Reproduce:
pip install google-adk==2.11.0BaseLlmthat calls both tools in parallel).create_ticketreturns at once;notify_oncallraises 200 ms later. The script prints what the session kept.Expected Behavior:
The session keeps
create_ticket's real response, so the next turn shows the model that the ticket exists. Thefailing call gets no response, and nothing is made up. ADK already does this in two similar places: the confirmation
pause (#6732, 67d3e49) and workflow nodes (2c759e9, "keep outputs of sibling nodes that finished in the same tick as
a failing node").
Observed Behavior:
The ticket was created, but the session has no response for it. On the next turn the call is dropped
(
Dropping function calls with no matching function response: [...]), so the model has no record of it.Environment Details:
mainat fac77beModel Information:
BaseLlm); also seen through LiteLlm, see below.馃煛 Optional Information
Regression: Not checked.
Additional Context:
LiteLlm, asked "Please finish thetask.", it re-created the committed record in 10/10 runs on 2.11.0, and 0/10 with a fix.
stripping orphaned calls instead of making up responses; the fix I have keeps only the tools' real results and makes
nothing up.
that finished, and once the agent has stopped on the error they're persisted before the error event; the error
itself is unchanged.
Minimal Reproduction Code:
How often has this issue occurred?: