Skip to content

Gemini(use_interactions_api=True) streaming does not coalesce consecutive text/thought delta Parts #7468

Description

@Dumeng

🔴 Required Information

Please ensure all items in this section are completed to allow for efficient
triaging. Requests without complete information may be rejected / deprioritized.
If an item is not applicable to you - please mark it as N/A

Describe the Bug:
When Gemini(use_interactions_api=True) is used with streaming (stream=True / RunConfig(streaming_mode=StreamingMode.SSE)), the final non-partial LlmResponse (partial=False, turn_complete=True) emitted on interaction.complete does not coalesce consecutive text or thought delta chunks. Instead, LlmResponse.content.parts contains one types.Part per SSE delta event (often 40–100+ Part(text=...) entries for a single response).
By contrast, on the standard generate_content_stream path (use_interactions_api=False), StreamingResponseAggregator._flush_text_buffer_to_sequence (src/google/adk/utils/streaming_utils.py) concatenates consecutive text deltas (''.join(self._current_text_buffer)) into a single types.Part(text=buffered_text) and consecutive thought deltas into a single types.Part(text=buffered_text, thought=True) before emitting the final partial=False LlmResponse.

Steps to Reproduce:
Please provide a numbered list of steps to reproduce the behavior:

  1. Install google-adk==2.11.0.
  2. Configure an LlmAgent with Gemini(model="gemini-3-flash-preview", use_interactions_api=True) and run it with streaming enabled (RunConfig(streaming_mode=StreamingMode.SSE)), or run the standalone reproduction script in Minimal Reproduction Code below.
  3. Inspect the final non-partial Event / LlmResponse (partial=False, turn_complete=True).
  4. Observe that final_response.content.parts contains dozens of individual token-delta types.Part(text=...) objects rather than a single coalesced text Part.

Expected Behavior:
On the final partial=False LlmResponse emitted when InteractionCompletedEvent (event_type == "interaction.complete") arrives, consecutive types.Part(text=..., thought=None/False) entries in _StreamState.parts should be merged into a single types.Part(text=...), and consecutive types.Part(text=..., thought=True) entries should be merged into a single types.Part(text=..., thought=True) (preserving thought_signature), matching the behavior of StreamingResponseAggregator in src/google/adk/utils/streaming_utils.py.

Observed Behavior:
In src/google/adk/models/interactions_utils.py:

  • _handle_text and _handle_thought_summary append a new types.Part directly to state.parts on every content.delta SSE event (TextDelta / ThoughtSummaryDelta).
  • When InteractionCompletedEvent is processed in convert_interaction_event_to_llm_response, it returns LlmResponse(content=types.Content(role='model', parts=state.parts), partial=False, turn_complete=True) without coalescing consecutive text or thought parts.
  • As a result:
    1. Persisted session events (Session.events) store 40–100+ fine-grained Part objects per turn, breaking downstream UIs and consumers that iterate event.content.parts expecting semantic content blocks.
    2. ADK's own AgentTool.run_async (src/google/adk/tools/agent_tool.py) joins last_content.parts using '\n'.join(_part_to_text(p) for p in last_content.parts if not p.thought), which corrupts sub-agent output by inserting a newline between every streamed token delta.

Environment Details:

  • ADK Library Version (pip show google-adk): 2.11.0
  • Desktop OS: Linux
  • Python Version (python -V): Python 3.14 (also reproduces on Python 3.12 / 3.13)

Model Information:

  • Are you using LiteLLM: No
  • Which model is being used: gemini-3.8-flash

🟡 Optional Information

Providing this information greatly speeds up the resolution process.

Regression:
No

Logs:
Please attach relevant logs. Wrap them in code blocks (```) or attach a
text file.

# Example final persisted Event.content.parts from a single streamed turn:
parts=[
  Part(thought=True, text='Thinking Process:\n\n1.  **Understand the Goal**: '),
  Part(thought=True, text='The user is asking for...'),
  Part(text='Based on the '),
  Part(text='data available, here '),
  Part(text='is the summary '),
  Part(text='from `'),
  Part(text='2026-03'),
  Part(text='-01` to `2026-03-31`...'),
]

Screenshots / Video:
If applicable, add screenshots or screen recordings to help explain
your problem.
N/A

Additional Context:
Add any other context about the problem here.

In src/google/adk/models/interactions_utils.py:

  • _handle_text (around line 883) does state.parts.append(part) for every TextDelta.
  • _handle_thought_summary (around line 991) does state.parts.append(part) for every ThoughtSummaryDelta.
  • convert_interaction_event_to_llm_response (around line 1272) handles InteractionCompletedEvent by passing parts=state.parts directly into types.Content(role='model', parts=state.parts) without merging adjacent text/thought parts.
    Coalescing adjacent text parts (and adjacent thought parts while preserving thought_signature) before constructing the final LlmResponse on InteractionCompletedEvent resolves the discrepancy with StreamingResponseAggregator.

Minimal Reproduction Code:
Please provide a code snippet or a link to a Gist/repo that isolates the issue.

from google.adk.models.interactions_utils import (
    _StreamState,
    convert_interaction_event_to_llm_response,
)
from google.genai import _interactions
state = _StreamState()
# Simulate streamed thought deltas followed by text deltas
for i, chunk in enumerate(["Thinking step 1. ", "Thinking step 2."]):
    thought_event = _interactions.types.InteractionContentDeltaEvent(
        event_type="content.delta",
        index=0,
        delta=_interactions.types.ThoughtSummaryDelta(
            type="thought_summary",
            content=_interactions.types.TextContent(type="text", text=chunk),
        ),
        event_id=f"thought_{i}",
    )
    convert_interaction_event_to_llm_response(thought_event, state)
for i, chunk in enumerate(["Hello", ", ", "world", "!"]):
    text_event = _interactions.types.InteractionContentDeltaEvent(
        event_type="content.delta",
        index=1,
        delta=_interactions.types.TextDelta(type="text", text=chunk),
        event_id=f"text_{i}",
    )
    convert_interaction_event_to_llm_response(text_event, state)
completed_event = _interactions.types.InteractionCompletedEvent(
    event_type="interaction.complete",
    interaction=_interactions.types.Interaction(id="int_123", status="completed"),
)
final_resp = convert_interaction_event_to_llm_response(completed_event, state)
print(f"Number of parts: {len(final_resp.content.parts)}")
print(final_resp.content.parts)
# Observed: 6 parts (2 thought parts + 4 text parts)
# Expected: 2 parts (1 coalesced thought part + 1 coalesced text part)
assert len(final_resp.content.parts) == 2
assert final_resp.content.parts[0].thought is True
assert final_resp.content.parts[0].text == "Thinking step 1. Thinking step 2."
assert final_resp.content.parts[1].text == "Hello, world!"

How often has this issue occurred?:

Always (100%)

Activity

  1. 13pgpg commented on Oct 9, 2026

    @13pgpg

    I'd like to work on this. I'll add a focused regression test for a completed Interactions API stream with adjacent text and thought deltas, then coalesce only adjacent compatible parts while preserving thought signatures and non-text boundaries. Could a maintainer assign this issue to @13pgpg if that scope looks right?

  2. self-assigned this
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions