Skip to content

Unintended user message injection breaks tool calling with LiteLLM + OpenAI/Azure #4249

Description

@GitMarco27

Describe the bug

When using LiteLLM with OpenAI / AzureOpenAI models, after tool calls, an user message "Handle the requests as specified in the System Instruction." is injected into the conversation history. This message triggers OpenAI's prompt injection safety guards and causes the request to fail.

openai.BadRequestError: Error code: 400 - {'error': {'message': "The response was filtered due to the prompt triggering Azure OpenAI's content management policy. Please modify your prompt and retry. To learn more about our content filtering policies please read our documentation: https://go.microsoft.com/fwlink/?linkid=2198766", 'type': None, 'param': 'prompt', 'code': 'content_filter', 'status': 400, 'innererror': {'code': 'ResponsibleAIPolicyViolation', 'content_filter_result': {'hate': {'filtered': False, 'severity': 'safe'}, 'jailbreak': {'filtered': True, 'detected': True}, 'self_harm': {'filtered': False, 'severity': 'safe'}, 'sexual': {'filtered': False, 'severity': 'safe'}, 'violence': {'filtered': False, 'severity': 'safe'}}}}}

Root Cause Analysis

The bug originates from an interaction between two functions in lite_llm.py:

  1. _part_has_payload() does not consider function_response as a valid payload for the model:
def _part_has_payload(part: types.Part) -> bool:
  """Checks whether a Part contains usable payload for the model."""
  if part.text:
    return True
  if part.inline_data and part.inline_data.data:
    return True
  if part.file_data and (part.file_data.file_uri or part.file_data.data):
    return True
  return False
  1. _append_fallback_user_content_if_missing() iterates through llm_request.contents looking for a user Content with payload. Tool responses (function_response) are represented as Content with role='user'. Since _part_has_payload() returns False for function_response parts, the function incorrectly appends a fallback text part.
def _append_fallback_user_content_if_missing(
    llm_request: LlmRequest,
) -> None:
  """Ensures there is a user message with content for LiteLLM backends.

  Args:
    llm_request: The request that may need a fallback user message.
  """
  for content in reversed(llm_request.contents):
    if content.role == "user":
      parts = content.parts or []
      if any(_part_has_payload(part) for part in parts):
        return
      if not parts:
        content.parts = []
      content.parts.append(
          types.Part.from_text(
              text="Handle the requests as specified in the System Instruction."
          )
      )
      return
  llm_request.contents.append(
      types.Content(
          role="user",
          parts=[
              types.Part.from_text(
                  text=(
                      "Handle the requests as specified in the System"
                      " Instruction."
                  )
              ),
          ],
      )
  )
  1. _content_to_message_param() then processes this Content. The newly added text part becomes a non_tool_part, which triggers the branch:
if tool_messages and non_tool_parts:
    follow_up = await _content_to_message_param(
        types.Content(role=content.role, parts=non_tool_parts),
        provider=provider,
    )
    follow_up_messages = (
        follow_up if isinstance(follow_up, list) else [follow_up]
    )
    return tool_messages + follow_up_messages

This creates an extra user message with the fallback text "Handle the requests as specified in the System Instruction.", which OpenAI's safety systems flag as potential prompt injection.

Previous Behavior

In previous versions, _content_to_message_param handled this by returning early when tool_messages existed, so the new user part was never really provided to the model invocation:

if tool_messages:
    return tool_messages if len(tool_messages) > 1 else tool_messages[0]

This prevented the fallback text from being converted into a follow-up user message.


To Reproduce

  1. Install google-adk>1.22.1
  2. Create an agent with at least one tool using LiteLLM with an OpenAI model
  3. Trigger a conversation that results in a tool call
  4. Observe the LLM request includes an extra user message: "Handle the requests as specified in the System Instruction."
  5. With OpenAI models, this triggers prompt injection safety guards

Minimal reproduction code

import asyncio
import os

from dotenv import load_dotenv
from google.adk.agents import Agent
from google.adk.models.lite_llm import LiteLlm
from google.adk.runners import Runner
from google.adk.sessions import InMemorySessionService
from google.genai import types
from google.genai.types import GenerateContentConfig

load_dotenv()


def get_litellm_model() -> LiteLlm:
    deployment_id = os.getenv("AZURE_OPENAI_DEPLOYMENT_ID")
    return LiteLlm(
        model=f"azure/{deployment_id}",
        stream=True,
    )


async def add(a: float, b: float) -> float:
    return a + b


def create_simple_agent() -> Agent:
    model = get_litellm_model()

    instructions = """
    You are a helpful mathematical assistant with access to a calculator tool.
    """

    agent = Agent(
        name="Claudia",
        model=model,
        generate_content_config=GenerateContentConfig(temperature=0.0),
        instruction=instructions,
        description="A simple agent that can perform basic mathematical calculations",
        tools=[add],
    )

    return agent


async def run_agent_single_query(query: str):
    agent = create_simple_agent()
    session_service = InMemorySessionService()

    runner = Runner(
        agent=agent,
        app_name="SimpleADKApp",
        session_service=session_service,
        auto_create_session=True,
    )

    user_id = "user_123"
    session_id = "session_001"

    content = types.Content(role="user", parts=[types.Part(text=query)])

    response_text = ""
    async for event in runner.run_async(
        user_id=user_id,
        session_id=session_id,
        new_message=content,
    ):
        if event.content and event.content.parts:
            response_text = event.content.parts[0].text or ""

    print(f"Agent Response:\n{response_text}\n")
    return response_text


async def main():
    await run_agent_single_query("Ciao! Quanto fa 3 + 3?")


if __name__ == "__main__":
    asyncio.run(main())

Error / Stacktrace

litellm.exceptions.ContentPolicyViolationError: litellm.BadRequestError: litellm.ContentPolicyViolationError: litellm.ContentPolicyViolationError: AzureException - The response was filtered due to the prompt triggering Azure OpenAI's content management policy. Please modify your prompt and retry. To learn more about our content filtering policies please read our documentation: https://go.microsoft.com/fwlink/?linkid=2198766

Expected behavior

Content objects containing only function_response parts should NOT have the fallback text appended. The _part_has_payload() function should return True for parts that have function_response, recognizing that tool responses are valid payload.


Proposed fix

  1. Add a function_response check in _part_has_payload() in order to avoid appending this fallback for function responses:
def _part_has_payload(part: types.Part) -> bool:
  """Checks whether a Part contains usable payload for the model."""
  if part.text:
    return True
  if part.inline_data and part.inline_data.data:
    return True
  if part.file_data and (part.file_data.file_uri or part.file_data.data):
    return True
  if part.function_response:
    return True
  return False
  1. Remove _append_fallback_user_content_if_missing: avoid adding a fallback user message at all as this is unexpected, dangerous and really similar to a silent failure.

Desktop

  • OS: Ubuntu 20.04.6 LTS
  • Python version: 3.12.11
  • ADK version: > 1.22.1

Model Information

  • Are you using LiteLLM: Yes
  • Which model is being used: OpenAI models via LiteLLM (gpt-4.1)

Additional context

  1. The bug specifically affects the flow:

    • User sends message -> Model responds with tool call -> Tool executes -> Tool response (function_response) is added to history -> _append_fallback_user_content_if_missing incorrectly modifies this Content -> Extra user message is created
  2. This issue may not manifest with Gemini models accessed directly (without LiteLLM) because they may handle the message format differently.

  3. This bug forbids the use of any tool calls in conjunction with LiteLLM and OpenAI models and so does not allow me to use the latest releases of Google ADK

Activity

  1. added
    models[Component] This issue is related to model support
    on Jan 23, 2026
  2. added theissue type on Jan 23, 2026
  3. sandangel commented on Jan 26, 2026

    @sandangel

    I have the same issue. Is there a workaround for this issue?

  4. GitMarco27 commented on Jan 26, 2026

    @GitMarco27
    ContributorAuthor

    I have the same issue. Is there a workaround for this issue?

    Actually solved by freezing the library version to google-adk==1.22.1. A workaround is feasible, but requires a ton of overriding as the issue is nested into the adk codebase.

  5. sandangel commented on Jan 26, 2026

    @sandangel
    class OpenAILiteLlm(LiteLlm):
        """LiteLlm with fix for OpenAI API compatibility.
    
        ADK's _append_fallback_user_content_if_missing adds an extra user message
        after tool responses, which breaks OpenAI's API format. This class prevents
        that by ensuring function_response parts have payload (inline_data) so the
        fallback check passes.
        """
    
        async def generate_content_async(
            self, llm_request: LlmRequest, stream: bool = False
        ) -> AsyncGenerator[LlmResponse, None]:
            # Add inline_data to function_response parts so _part_has_payload returns
            # True and _append_fallback_user_content_if_missing doesn't add fallback
            for content in llm_request.contents or []:
                if content.role != "user":
                    continue
                for part in content.parts or []:
                    if part.function_response and not part.inline_data:
                        part.inline_data = types.Blob(data=b" ", mime_type="text/plain")
    
            async for response in super().generate_content_async(llm_request, stream):
                yield response

    I ended up with this fix, tricking ADK to think this is a message with user payload so it won't auto append the user message again.

  6. GitMarco27 commented on Jan 26, 2026

    @GitMarco27
    ContributorAuthor

    I ended up with this fix, tricking ADK to think this is a message with user payload so it won't auto append the user message again.

    A bit dirty, but does the job!

  7. self-assigned this
    on Jan 26, 2026
  8. bhakta0007 commented on Jan 27, 2026

    @bhakta0007

    +1 Bump.

    I hit this issue and moved back one version - only to hit another issue with compaction.

  9. CactionCCY commented on Jan 27, 2026

    @CactionCCY

    Same Issue as well, triggering Azure JailBreak detection during tool calling.

  10. seandi commented on Jan 28, 2026

    @seandi

    I am also facing the same issue!

  11. dineshkrishna9999 commented on Jan 28, 2026

    @dineshkrishna9999

    I was able to debug and reproduce this issue.
    Normally, when a tool returns a response, the function_response should be sent back to the model so the model can generate the final user-facing reply.

    However, in lite_llm.py, ADK only checks for payload in Part(text) (and a few other fields) and does not treat Part(function_response) as valid payload. Because of this, the payload check returns false even though a valid tool response exists.

    As a result, the following fallback logic is triggered, which appends the default text:
    "Handle the requests as specified in the System Instruction."
    This injected text is interpreted by Azure OpenAI as prompt injection / jailbreak content, which causes the request to fail with a content policy violation.
    I was able to recreate the entire scenario locally. As a workaround, treating function_response as valid payload fixes the issue:

    def _part_has_payload(part: types.Part) -> bool:
        """Checks whether a Part contains usable payload for the model."""
        if part.text:
            return True
        if part.function_response:  # added
            return True
        if part.inline_data and part.inline_data.data:
            return True
        if part.file_data and (part.file_data.file_uri or part.file_data.data):
            return True
        return False
    

    With this change, the fallback text is not injected and the model proceeds correctly after the tool response.

  12. KrshnKush commented on Jan 29, 2026

    @KrshnKush
    Contributor

    @GitMarco27 @dineshkrishna9999 @seandi @CactionCCY @bhakta0007 @sandangel

    https://github.com/google/adk-python/blob/main/src/google/adk/models/lite_llm.py#L503

    This statement is the culprit: Handle the requests as specified in the System Instruction.

    Instead use: Handle the incoming request according to the provided requirements.

    Image
  13. added a commit that references this issue on Jan 29, 2026
    92e09dd
  14. GitMarco27 commented on Jan 29, 2026

    @GitMarco27
    ContributorAuthor

    Handle the incoming request according to the provided requirements.

    I think that this instruction might still be interpreted as a jailbreak or an injection. Overall, it doesn't solve the actual issue, the introduction of unintended and unexpected context into the LLM request.

  15. KrshnKush commented on Jan 29, 2026

    @KrshnKush
    Contributor

    Handle the incoming request according to the provided requirements.

    I think that this instruction might still be interpreted as a jailbreak or an injection. Overall, it doesn't solve the actual issue, the introduction of unintended and unexpected context into the LLM request.

    @GitMarco27 It worked fine, check the second pull request mentioned above, attached screenshots.

  16. GitMarco27 commented on Jan 29, 2026

    @GitMarco27
    ContributorAuthor

    It worked fine, check the second pull request mentioned above, attached screenshots.

    I see, but looks like a workaround. It might break again as soon as Openai / Azure Openai updates its filters' models and policies.

  17. KrshnKush commented on Jan 29, 2026

    @KrshnKush
    Contributor

    It worked fine, check the second pull request mentioned above, attached screenshots.

    I see, but looks like a workaround. It might break again as soon as Openai / Azure Openai updates its filters' models and policies.

    Maybe.

  18. added a commit that references this issue on Jan 30, 2026
    d0102ec
  19. added a commit that references this issue on Aug 11, 2026
    012e44c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

models[Component] This issue is related to model support

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions