Repository navigation
Conversation
_build_rate_limit_error() interpolates the provider detail verbatim, so
a quota error still reaches the modeller as the provider's raw JSON
report. For a Gemini free-tier 429 the rewritten message is 1308
characters over 36 lines, longer than the original error, while the
retryDelay the provider supplies is never shown.
Extract the human-readable message and the retry delay from the payload
and drop the rest, so the error reads as one line, and use the reported
delay in the guidance ("Please wait 27s") instead of the generic "a few
minutes". Payloads without a message, or that do not parse, are
collapsed to a single line and truncated, which leaves short plain-text
details unchanged. Repeated "litellm.RateLimitError:" prefixes are
stripped rather than the first one only.
Gemini free-tier 429: 1308 chars / 36 lines -> 317 chars / 1 line.
Fixes mesa#257
Contributor
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pre-PR Checklist
Summary
A quota error still reaches the modeller as the provider's raw JSON report.
_build_rate_limit_error()wraps that payload in friendly text instead of replacing it, so the result is longer than the original error. This PR condenses the payload to one readable line and surfaces the retry delay the provider already reports.Bug / Issue
Fixes #257 ("200+ lines of unreadable JSON").
_build_rate_limit_error()(module_llm.py:160) adds a headline, guidance and a docs link, then interpolateserror.messageverbatim. For a Gemini free-tier 429 the payload is a nestedQuotaFailure/Help/RetryInforeport, so the "friendly" message measured onmainis 1308 characters over 36 lines, against 1077 characters for the raw error.Two further gaps:
retryDelay: "27s", which is exactly the actionable part, and it is dropped. The user is told to "wait a few minutes".litellm.RateLimitError:prefix is stripped, and LiteLLM often nests it twice.Implementation
_condense_provider_detail()locates the JSON object in the detail, extracts the human-readablemessageand anyretryDelay/retry_delay/retry-afterat any depth, and drops the rest.test_generate_rewrites_rate_limit_error_with_*intact.litellm.RateLimitError:prefixes are stripped.Result for the Gemini free-tier payload:
(one line in the terminal; wrapped here for readability)
1308 chars / 36 lines → 317 chars / 1 line. The provider's own status, code and quota identifiers stay available on the exception object; only the message shown to the user is condensed.
Alternatives rejected: truncating the detail blindly (cuts the sentence that matters and keeps the JSON noise); logging the full payload at DEBUG (a second knob for the same problem, and the raw exception is still chained via
from error).Testing
New
TestRateLimitErrorReadability(6 tests), 5 of which fail onmain:quotaMetric/@type.litellm.RateLimitError:prefixes are stripped.agenerate()condenses identically.Full suite: 1558 passed.
pre-commit runpasses.Additional Notes
RateLimitErroris retryable and the retry never ends. My PR for Issues with Gemini quickstart flow (retry behavior, model errors, and tool usage #266 fixes that; this PR is what the user finally sees. Either can merge first.module_llm.pyat the same anchor points, so whichever merges second needs a trivial rebase. I will rebase immediately on request.