🔴 Required Information
Please ensure all items in this section are completed to allow for efficient
triaging. Requests without complete information may be rejected / deprioritized.
If an item is not applicable to you - please mark it as N/A
Describe the Bug:
When Gemini(use_interactions_api=True) is used with streaming (stream=True / RunConfig(streaming_mode=StreamingMode.SSE)), the final non-partial LlmResponse (partial=False, turn_complete=True) emitted on interaction.complete does not coalesce consecutive text or thought delta chunks. Instead, LlmResponse.content.parts contains one types.Part per SSE delta event (often 40–100+ Part(text=...) entries for a single response).
By contrast, on the standard generate_content_stream path (use_interactions_api=False), StreamingResponseAggregator._flush_text_buffer_to_sequence (src/google/adk/utils/streaming_utils.py) concatenates consecutive text deltas (''.join(self._current_text_buffer)) into a single types.Part(text=buffered_text) and consecutive thought deltas into a single types.Part(text=buffered_text, thought=True) before emitting the final partial=False LlmResponse.
Steps to Reproduce:
Please provide a numbered list of steps to reproduce the behavior:
- Install
google-adk==2.11.0.
- Configure an
LlmAgent with Gemini(model="gemini-3-flash-preview", use_interactions_api=True) and run it with streaming enabled (RunConfig(streaming_mode=StreamingMode.SSE)), or run the standalone reproduction script in Minimal Reproduction Code below.
- Inspect the final non-partial
Event / LlmResponse (partial=False, turn_complete=True).
- Observe that
final_response.content.parts contains dozens of individual token-delta types.Part(text=...) objects rather than a single coalesced text Part.
Expected Behavior:
On the final partial=False LlmResponse emitted when InteractionCompletedEvent (event_type == "interaction.complete") arrives, consecutive types.Part(text=..., thought=None/False) entries in _StreamState.parts should be merged into a single types.Part(text=...), and consecutive types.Part(text=..., thought=True) entries should be merged into a single types.Part(text=..., thought=True) (preserving thought_signature), matching the behavior of StreamingResponseAggregator in src/google/adk/utils/streaming_utils.py.
Observed Behavior:
In src/google/adk/models/interactions_utils.py:
_handle_text and _handle_thought_summary append a new types.Part directly to state.parts on every content.delta SSE event (TextDelta / ThoughtSummaryDelta).
- When
InteractionCompletedEvent is processed in convert_interaction_event_to_llm_response, it returns LlmResponse(content=types.Content(role='model', parts=state.parts), partial=False, turn_complete=True) without coalescing consecutive text or thought parts.
- As a result:
- Persisted session events (
Session.events) store 40–100+ fine-grained Part objects per turn, breaking downstream UIs and consumers that iterate event.content.parts expecting semantic content blocks.
- ADK's own
AgentTool.run_async (src/google/adk/tools/agent_tool.py) joins last_content.parts using '\n'.join(_part_to_text(p) for p in last_content.parts if not p.thought), which corrupts sub-agent output by inserting a newline between every streamed token delta.
Environment Details:
- ADK Library Version (pip show google-adk):
2.11.0
- Desktop OS:
Linux
- Python Version (python -V):
Python 3.14 (also reproduces on Python 3.12 / 3.13)
Model Information:
- Are you using LiteLLM: No
- Which model is being used:
gemini-3.8-flash
🟡 Optional Information
Providing this information greatly speeds up the resolution process.
Regression:
No
Logs:
Please attach relevant logs. Wrap them in code blocks (```) or attach a
text file.
# Example final persisted Event.content.parts from a single streamed turn:
parts=[
Part(thought=True, text='Thinking Process:\n\n1. **Understand the Goal**: '),
Part(thought=True, text='The user is asking for...'),
Part(text='Based on the '),
Part(text='data available, here '),
Part(text='is the summary '),
Part(text='from `'),
Part(text='2026-03'),
Part(text='-01` to `2026-03-31`...'),
]
Screenshots / Video:
If applicable, add screenshots or screen recordings to help explain
your problem.
N/A
Additional Context:
Add any other context about the problem here.
In src/google/adk/models/interactions_utils.py:
_handle_text (around line 883) does state.parts.append(part) for every TextDelta.
_handle_thought_summary (around line 991) does state.parts.append(part) for every ThoughtSummaryDelta.
convert_interaction_event_to_llm_response (around line 1272) handles InteractionCompletedEvent by passing parts=state.parts directly into types.Content(role='model', parts=state.parts) without merging adjacent text/thought parts.
Coalescing adjacent text parts (and adjacent thought parts while preserving thought_signature) before constructing the final LlmResponse on InteractionCompletedEvent resolves the discrepancy with StreamingResponseAggregator.
Minimal Reproduction Code:
Please provide a code snippet or a link to a Gist/repo that isolates the issue.
from google.adk.models.interactions_utils import (
_StreamState,
convert_interaction_event_to_llm_response,
)
from google.genai import _interactions
state = _StreamState()
# Simulate streamed thought deltas followed by text deltas
for i, chunk in enumerate(["Thinking step 1. ", "Thinking step 2."]):
thought_event = _interactions.types.InteractionContentDeltaEvent(
event_type="content.delta",
index=0,
delta=_interactions.types.ThoughtSummaryDelta(
type="thought_summary",
content=_interactions.types.TextContent(type="text", text=chunk),
),
event_id=f"thought_{i}",
)
convert_interaction_event_to_llm_response(thought_event, state)
for i, chunk in enumerate(["Hello", ", ", "world", "!"]):
text_event = _interactions.types.InteractionContentDeltaEvent(
event_type="content.delta",
index=1,
delta=_interactions.types.TextDelta(type="text", text=chunk),
event_id=f"text_{i}",
)
convert_interaction_event_to_llm_response(text_event, state)
completed_event = _interactions.types.InteractionCompletedEvent(
event_type="interaction.complete",
interaction=_interactions.types.Interaction(id="int_123", status="completed"),
)
final_resp = convert_interaction_event_to_llm_response(completed_event, state)
print(f"Number of parts: {len(final_resp.content.parts)}")
print(final_resp.content.parts)
# Observed: 6 parts (2 thought parts + 4 text parts)
# Expected: 2 parts (1 coalesced thought part + 1 coalesced text part)
assert len(final_resp.content.parts) == 2
assert final_resp.content.parts[0].thought is True
assert final_resp.content.parts[0].text == "Thinking step 1. Thinking step 2."
assert final_resp.content.parts[1].text == "Hello, world!"
How often has this issue occurred?:
Always (100%)
🔴 Required Information
Please ensure all items in this section are completed to allow for efficient
triaging. Requests without complete information may be rejected / deprioritized.
If an item is not applicable to you - please mark it as N/A
Describe the Bug:
When
Gemini(use_interactions_api=True)is used with streaming (stream=True/RunConfig(streaming_mode=StreamingMode.SSE)), the final non-partialLlmResponse(partial=False, turn_complete=True) emitted oninteraction.completedoes not coalesce consecutive text or thought delta chunks. Instead,LlmResponse.content.partscontains onetypes.Partper SSE delta event (often 40–100+Part(text=...)entries for a single response).By contrast, on the standard
generate_content_streampath (use_interactions_api=False),StreamingResponseAggregator._flush_text_buffer_to_sequence(src/google/adk/utils/streaming_utils.py) concatenates consecutive text deltas (''.join(self._current_text_buffer)) into a singletypes.Part(text=buffered_text)and consecutive thought deltas into a singletypes.Part(text=buffered_text, thought=True)before emitting the finalpartial=FalseLlmResponse.Steps to Reproduce:
Please provide a numbered list of steps to reproduce the behavior:
google-adk==2.11.0.LlmAgentwithGemini(model="gemini-3-flash-preview", use_interactions_api=True)and run it with streaming enabled (RunConfig(streaming_mode=StreamingMode.SSE)), or run the standalone reproduction script in Minimal Reproduction Code below.Event/LlmResponse(partial=False, turn_complete=True).final_response.content.partscontains dozens of individual token-deltatypes.Part(text=...)objects rather than a single coalesced textPart.Expected Behavior:
On the final
partial=FalseLlmResponseemitted whenInteractionCompletedEvent(event_type == "interaction.complete") arrives, consecutivetypes.Part(text=..., thought=None/False)entries in_StreamState.partsshould be merged into a singletypes.Part(text=...), and consecutivetypes.Part(text=..., thought=True)entries should be merged into a singletypes.Part(text=..., thought=True)(preservingthought_signature), matching the behavior ofStreamingResponseAggregatorinsrc/google/adk/utils/streaming_utils.py.Observed Behavior:
In
src/google/adk/models/interactions_utils.py:_handle_textand_handle_thought_summaryappend a newtypes.Partdirectly tostate.partson everycontent.deltaSSE event (TextDelta/ThoughtSummaryDelta).InteractionCompletedEventis processed inconvert_interaction_event_to_llm_response, it returnsLlmResponse(content=types.Content(role='model', parts=state.parts), partial=False, turn_complete=True)without coalescing consecutive text or thought parts.Session.events) store 40–100+ fine-grainedPartobjects per turn, breaking downstream UIs and consumers that iterateevent.content.partsexpecting semantic content blocks.AgentTool.run_async(src/google/adk/tools/agent_tool.py) joinslast_content.partsusing'\n'.join(_part_to_text(p) for p in last_content.parts if not p.thought), which corrupts sub-agent output by inserting a newline between every streamed token delta.Environment Details:
2.11.0LinuxPython 3.14(also reproduces onPython 3.12/3.13)Model Information:
gemini-3.8-flash🟡 Optional Information
Providing this information greatly speeds up the resolution process.
Regression:
No
Logs:
Please attach relevant logs. Wrap them in code blocks (```) or attach a
text file.
Screenshots / Video:
If applicable, add screenshots or screen recordings to help explain
your problem.
N/A
Additional Context:
Add any other context about the problem here.
In
src/google/adk/models/interactions_utils.py:_handle_text(around line 883) doesstate.parts.append(part)for everyTextDelta._handle_thought_summary(around line 991) doesstate.parts.append(part)for everyThoughtSummaryDelta.convert_interaction_event_to_llm_response(around line 1272) handlesInteractionCompletedEventby passingparts=state.partsdirectly intotypes.Content(role='model', parts=state.parts)without merging adjacent text/thought parts.Coalescing adjacent text parts (and adjacent thought parts while preserving
thought_signature) before constructing the finalLlmResponseonInteractionCompletedEventresolves the discrepancy withStreamingResponseAggregator.Minimal Reproduction Code:
Please provide a code snippet or a link to a Gist/repo that isolates the issue.
How often has this issue occurred?:
Always (100%)