Skip to content

LiteLlm streaming drops a thought_signature sent on a tool-call delta without the call's name or arguments聽#7485

Description

@vetler

馃敶 Required Information

Describe the Bug:
With LiteLlm and stream=True, a thought_signature that arrives on a tool-call delta of its own, without the call's name or arguments, is dropped. _model_response_to_chunk treats a delta with neither a name nor arguments as empty and skips it, so the signature never reaches the function-call part. Gemini 3 then rejects the next request in the turn ("Function call is missing a thought_signature in functionCall parts").

This is the remaining case from #7438. #7441 (landed as 3abea3e) keeps a signature that arrives on the same delta as the call's name, which is what Vertex AI sends today. The signature-only case was raised on #7438 and left for a follow-up.

Steps to Reproduce:

  1. Check out main (reproduced at f5db310).
  2. Run the script under Minimal Reproduction Code. It mocks the LiteLLM client, so no credentials are needed.
  3. The stream sends the call's name and arguments on one delta and the signature on a second delta for the same index.

Expected Behavior:

function calls: [('call_1', 'get_weather')]
thought_signature: b'late-signature'

Observed Behavior:

function calls: [('call_1', 'get_weather')]
thought_signature: None

Environment Details:

  • ADK Library Version (pip show google-adk): main at f5db310
  • Desktop OS: macOS
  • Python Version (python -V): 3.11.16

Model Information:

  • Are you using LiteLLM: Yes (litellm 1.85.7)
  • Which model is being used: Gemini 3 (gemini-3-flash-preview)

馃煛 Optional Information

Regression:
No. Before #7441 every streamed signature was lost; this is the one shape #7441 does not cover.

Additional Context:
I have not seen this shape from a real endpoint. In live runs with gemini-3-flash-preview, both Vertex AI's OpenAI-compatible endpoint and LiteLLM's vertex_ai provider put the signature on the delta that names the call (raw deltas in #7441). So nobody should be hitting this today; it matters for a provider or a LiteLLM version that splits the signature into its own delta.

A fix needs to avoid opening a new call for such a delta, which would reach the model as a call with no name, and to attach the signature by the delta's own index, since the fallback index used for providers with improper indexing may already point past that call.

Fix: #7467.

Minimal Reproduction Code:

"""LiteLlm drops a thought_signature that streams on a tool-call delta of its own.

No network or credentials: the LiteLLM client is mocked. The first delta
carries the call's name and arguments, a second delta for the same index
carries only the signature, then the stream finishes.
"""

import asyncio
import base64
from unittest.mock import AsyncMock, MagicMock

import litellm
from google.adk.models.lite_llm import LiteLlm
from google.adk.models.llm_request import LlmRequest
from google.genai import types

SIGNATURE = base64.b64encode(b"late-signature").decode()


def delta(tool_call, finish_reason=None):
  return litellm.ModelResponseStream(choices=[{
      "index": 0,
      "finish_reason": finish_reason,
      "delta": {"role": "assistant", "tool_calls": [tool_call]},
  }])


async def stream():
  yield delta({
      "index": 0,
      "id": "call_1",
      "type": "function",
      "function": {"name": "get_weather", "arguments": '{"city": "Oslo"}'},
  })
  # Same call, signature only: no name, no arguments.
  yield delta({
      "index": 0,
      "type": "function",
      "function": {"name": None, "arguments": ""},
      "extra_content": {"google": {"thought_signature": SIGNATURE}},
  })
  yield litellm.ModelResponseStream(
      choices=[{"index": 0, "finish_reason": "tool_calls", "delta": {}}]
  )


async def main():
  llm = LiteLlm(model="google/gemini-3-flash-preview", custom_llm_provider="openai")
  llm.llm_client = MagicMock()
  llm.llm_client.acompletion = AsyncMock(return_value=stream())
  request = LlmRequest(
      model="google/gemini-3-flash-preview",
      contents=[types.Content(role="user", parts=[types.Part(text="Weather in Oslo?")])],
      config=types.GenerateContentConfig(),
  )
  final = None
  async for response in llm.generate_content_async(request, stream=True):
    if not response.partial:
      final = response
  calls = [p for p in final.content.parts if p.function_call]
  print("function calls:", [(p.function_call.id, p.function_call.name) for p in calls])
  print("thought_signature:", calls[0].thought_signature)


asyncio.run(main())

How often has this issue occurred?:

  • Always (100%), for this delta shape

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions