Skip to content

Python: Honor the configured model in streamed Gemini chunks - #9139

Open
Ayoub (Ayoubhm07) wants to merge 3 commits into
microsoft:mainfrom
Ayoubhm07:fix/gemini-stream-model-fallback
Open

Ayoub (Ayoubhm07) wants to merge 3 commits into
microsoft:mainfrom
Ayoubhm07:fix/gemini-stream-model-fallback

Conversation

@Ayoubhm07

Copy link
Copy Markdown

Motivation & Context

_process_chunk passes chunk.model_version straight into ChatResponseUpdate.model, while _process_generate_response falls back to self.model for the same field 33 lines above. _types.py:2229-2230 keeps the last non-None update value and the Gemini client has no catch-up afterwards, so a stream whose chunks carry no model_version ends up with ChatResponse.model is None where the non-streaming path reports the model.

observability.py:3551-3552 sets gen_ai.response.model from that field, and RESPONSE_MODEL is in GEN_AI_METRIC_ATTRIBUTES (:3558): _filter_metric_attributes drops absent keys before _capture_response records the token and duration histograms (:3604-3610). On Gemini streams those lose their per-model dimension, and _backfill_request_model (:1905) cannot recover the span name either.

Description & Review Guide

  • What are the major changes? One line, applying the expression already on :1058 to :1091. Both methods take the same types.GenerateContentResponse (:1034-1036, :1066), and model_version is Optional[str] in the pinned google-genai 2.27.0 the SDK documents per-chunk absence for the very next field of that class, prompt_feedback: "Sent only in the first stream chunk."
  • What is the impact of these changes? A chunk that carries model_version is untouched. Only the update's model changes, and through aggregation ChatResponse.model. Two tests added next to the existing non-streaming one; only the fallback test fails on main, the other guards the pass-through and passes already. pytest packages/gemini/tests 254 passed 10 skipped, ruff check and ruff format --check clean, pyright 0 errors on the changed source, mypy packages/gemini/tests clean all at the versions in uv.lock.
  • What do you want reviewers to focus on? Whether self.model is the right value here: it is the client default, not the model resolved from options at :549, which _process_chunk does not receive. That matches :1058 exactly, so I kept parity rather than changing both happy to thread the resolved model into the two parsers instead if you prefer. MistralChatClient has the same asymmetry at :800 / :830, though its SDK types CompletionChunk.model as required so the gap is not reachable there; happy to follow up separately. ollama has no fallback on either path, so it is self-consistent and out of scope.

Related Issue

Fixes #9138

Contribution Checklist

  • The code builds clean without any errors or warnings
  • All unit tests pass, and I have added new tests where possible
  • The PR follows the Contribution Guidelines
  • This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
  • This is not a breaking change.

_process_chunk passed chunk.model_version straight into
ChatResponseUpdate.model, while _process_generate_response falls back to
self.model for the same field 33 lines above. Both take the same
types.GenerateContentResponse, and model_version is Optional[str] in the
pinned google-genai.

ChatResponse.model is built only from the updates, so a stream whose
chunks carry no model_version ended up with model=None. The
observability layer reads that field for gen_ai.response.model, which is
a metric dimension of the token and duration histograms, so those lost
their per-model breakdown on streams.

Apply the same expression the non-streaming path already uses. A chunk
that carries model_version is untouched.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The fallback mislabels responses when request options override the client’s default model.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Adds Gemini streaming model fallback to preserve response metadata and observability dimensions.

Changes:

  • Falls back to the client model when streamed chunks omit model_version.
  • Adds tests for pass-through and fallback behavior.
File Description
_chat_client.py Adds streamed model fallback.
test_gemini_client.py Tests streamed model handling.

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/packages/gemini/agent_framework_gemini/_chat_client.py Outdated
_process_chunk fell back to self.model, which is the client default. A
call that overrides the model through options was then labeled with the
client default instead of the model that actually ran, including in the
observability metrics.

_prepare_request already resolves the model and it was in scope at the
call site, so _process_chunk now takes it and uses it as the fallback.
Adds a streaming test with an options override.
@Ayoubhm07
Ayoub (Ayoubhm07) deployed to github-app-auth October 7, 2026 09:29 — with GitHub Actions Active
@Ayoubhm07

Ayoub (Ayoubhm07) commented Oct 7, 2026 •

Copy link
Copy Markdown
Author

_prepare_request already resolves the model and it was in scope at the call site so _process_chunk now takes it and uses it as the fallback instead of self.model

I added test_get_response_streaming_model_falls_back_to_requested_model: a gemini-2.5-flash client called with options={"model": "gemini-2.5-pro"} over a chunk with no model_version now reports gemini-2.5-pro.

_process_generate_response has the same client-default fallback at :1058. I left it alone to keep this PR on the streaming path happy to align it here or in a follow-up, whichever you prefer

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Oct 8, 2026
…odel-fallback

# Conflicts:
#	python/packages/gemini/agent_framework_gemini/_chat_client.py
@Ayoubhm07
Ayoub (Ayoubhm07) deployed to github-app-auth October 9, 2026 10:55 — with GitHub Actions Active
@Ayoubhm07

Copy link
Copy Markdown
Author

Resolved the conflict with main. The only overlap was the _process_chunk call, which 7a262d6a moved into the new try/finally that closes the SDK stream; the call now passes model from inside that block. The diff against main is unchanged otherwise: 4 lines in _chat_client.py and the 48 test lines.

This branch was successfully deployed

1 active deployment
github-app-auth — dd42966c Deployed Oct 9, 2026 by Ayoubhm07 via add_label #24872
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: [Bug]: Gemini streaming drops the model when a chunk has no model_version

3 participants