AzureCompletion._process_streaming_update accumulates streamed tool calls keyed on the position of each tool_call inside choice.delta.tool_calls (enumerate(...)), instead of on the index field of the tool call itself.
Azure AI Inference is OpenAI-compatible: in streaming mode it emits one tool_call entry per chunk, and that entry carries its own index field identifying which call the delta belongs to. With parallel tool calls, chunks for the second call arrive with index: 1 but sit at position 0 within their own chunk. So enumerate yields 0 for every chunk and all deltas collapse into a single accumulator slot.
The executor therefore receives one tool call instead of N, with the id of the first call, the name of the last, and the argument fragments concatenated in arrival order (usually invalid JSON).
Steps to Reproduce
from unittest.mock import patch
from azure.ai.inference.models import StreamingChatCompletionsUpdate
from crewai.llms.providers.azure.completion import AzureCompletion
def chunk(index, call_id, name, arguments):
return StreamingChatCompletionsUpdate(
id="chatcmpl-1", model="gpt-4o", created=None,
choices=[{
"index": 0,
"delta": {"role": "assistant", "content": None, "tool_calls": [{
"index": index, "id": call_id, "type": "function",
"function": {"name": name, "arguments": arguments},
}]},
"finish_reason": None,
}],
)
chunks = [
chunk(0, "call_aaa", "get_weather", '{"city":'),
chunk(0, "call_aaa", "get_weather", '"Paris"}'),
chunk(1, "call_bbb", "get_time", '{"tz":'),
chunk(1, "call_bbb", "get_time", '"UTC"}'),
]
llm = AzureCompletion(model="gpt-4o", api_key="k",
endpoint="https://test.openai.azure.com", stream=True)
with patch.object(llm._client, "complete") as m:
m.return_value = iter(chunks)
result = llm.call([{"role": "user", "content": "Weather and time in Paris?"}])
print(result)
Four chunks, alternating between the two parallel calls, exactly as the Azure streaming API delivers them.
Expected behavior
Two tool calls, each intact:
[
{"id": "call_aaa", "type": "function",
"function": {"name": "get_weather", "arguments": '{"city":"Paris"}'}},
{"id": "call_bbb", "type": "function",
"function": {"name": "get_time", "arguments": '{"tz":"UTC"}'}},
]
Screenshots/Code snippets
Actual output — one call, wrong name, and the two argument fragments concatenated into invalid JSON:
[{'id': 'call_aaa', 'type': 'function',
'function': {'name': 'get_time', 'arguments': '{"city":"Paris"}{"tz":"UTC"}'}}]
Root cause is lib/crewai/src/crewai/llms/providers/azure/completion.py around line 1000:
for idx, tool_call in enumerate(choice.delta.tool_calls):
if idx not in tool_calls:
tool_calls[idx] = {...}
For reference, the OpenAI provider in this repo already keys on the wire index (tool_call.index, e.g. lib/crewai/src/crewai/llms/providers/openai/completion.py (see the tool_index = tool_call.index aggregation)), so the two providers disagree on parallel tool calls.
Additional context
- The Azure
StreamingChatResponseToolCallUpdate model is a MutableMapping and does not declare index as an attribute (getattr(tc, "index") raises AttributeError), but the raw key survives deserialization, so tool_call.get("index") reads it correctly.
- Anything that issues two or more tool calls in one Azure streaming turn is affected; only the first call's id and only the last call's name survive, and the concatenated arguments fail
json.loads.
AzureCompletion._process_streaming_updateaccumulates streamed tool calls keyed on the position of eachtool_callinsidechoice.delta.tool_calls(enumerate(...)), instead of on theindexfield of the tool call itself.Azure AI Inference is OpenAI-compatible: in streaming mode it emits one
tool_callentry per chunk, and that entry carries its ownindexfield identifying which call the delta belongs to. With parallel tool calls, chunks for the second call arrive withindex: 1but sit at position0within their own chunk. Soenumerateyields0for every chunk and all deltas collapse into a single accumulator slot.The executor therefore receives one tool call instead of N, with the
idof the first call, thenameof the last, and the argument fragments concatenated in arrival order (usually invalid JSON).Steps to Reproduce
Four chunks, alternating between the two parallel calls, exactly as the Azure streaming API delivers them.
Expected behavior
Two tool calls, each intact:
[ {"id": "call_aaa", "type": "function", "function": {"name": "get_weather", "arguments": '{"city":"Paris"}'}}, {"id": "call_bbb", "type": "function", "function": {"name": "get_time", "arguments": '{"tz":"UTC"}'}}, ]Screenshots/Code snippets
Actual output — one call, wrong name, and the two argument fragments concatenated into invalid JSON:
[{'id': 'call_aaa', 'type': 'function', 'function': {'name': 'get_time', 'arguments': '{"city":"Paris"}{"tz":"UTC"}'}}]Root cause is
lib/crewai/src/crewai/llms/providers/azure/completion.pyaround line 1000:For reference, the OpenAI provider in this repo already keys on the wire index (
tool_call.index, e.g.lib/crewai/src/crewai/llms/providers/openai/completion.py (see thetool_index = tool_call.indexaggregation)), so the two providers disagree on parallel tool calls.Additional context
StreamingChatResponseToolCallUpdatemodel is aMutableMappingand does not declareindexas an attribute (getattr(tc, "index")raisesAttributeError), but the raw key survives deserialization, sotool_call.get("index")reads it correctly.json.loads.