Skip to content

Condense provider payloads in rate-limit errors - #341

Open
Lokkhita wants to merge 1 commit into
mesa:mainfrom
Lokkhita:fix/readable-rate-limit-error
Open

Lokkhita wants to merge 1 commit into
mesa:mainfrom
Lokkhita:fix/readable-rate-limit-error

Conversation

@Lokkhita

Copy link
Copy Markdown

Pre-PR Checklist

  • This PR is a bug fix, not a new feature or enhancement.

Summary

A quota error still reaches the modeller as the provider's raw JSON report. _build_rate_limit_error() wraps that payload in friendly text instead of replacing it, so the result is longer than the original error. This PR condenses the payload to one readable line and surfaces the retry delay the provider already reports.

Bug / Issue

Fixes #257 ("200+ lines of unreadable JSON").

_build_rate_limit_error() (module_llm.py:160) adds a headline, guidance and a docs link, then interpolates error.message verbatim. For a Gemini free-tier 429 the payload is a nested QuotaFailure / Help / RetryInfo report, so the "friendly" message measured on main is 1308 characters over 36 lines, against 1077 characters for the raw error.

Two further gaps:

  • The payload carries retryDelay: "27s", which is exactly the actionable part, and it is dropped. The user is told to "wait a few minutes".
  • Only the first litellm.RateLimitError: prefix is stripped, and LiteLLM often nests it twice.

Implementation

  • _condense_provider_detail() locates the JSON object in the detail, extracts the human-readable message and any retryDelay / retry_delay / retry-after at any depth, and drops the rest.
  • The guidance uses the reported delay when present: "Please wait 27s and try again", else the existing "a few minutes".
  • Fallback path: a payload that does not parse, or that carries no message, is collapsed to a single line and truncated to 200 characters with an ellipsis. Short plain-text details are therefore returned unchanged, which keeps the existing expectations in test_generate_rewrites_rate_limit_error_with_* intact.
  • Repeated litellm.RateLimitError: prefixes are stripped.

Result for the Gemini free-tier payload:

litellm.RateLimitError: Rate limit exceeded for model 'gemini/gemini-2.0-flash'.
VertexAIException - You exceeded your current quota, please check your plan and
billing details. Please wait 27s and try again, or switch to a different model.
To check your quota visit: https://ai.google.dev/gemini-api/docs/rate-limits

(one line in the terminal; wrapped here for readability)

1308 chars / 36 lines → 317 chars / 1 line. The provider's own status, code and quota identifiers stay available on the exception object; only the message shown to the user is condensed.

Alternatives rejected: truncating the detail blindly (cuts the sentence that matters and keeps the JSON noise); logging the full payload at DEBUG (a second knob for the same problem, and the raw exception is still chained via from error).

Testing

New TestRateLimitErrorReadability (6 tests), 5 of which fail on main:

  • A realistic Gemini quota payload renders as one line under 400 characters, keeps the quota sentence, and drops quotaMetric / @type.
  • The reported retry delay appears in the guidance, with the provider docs link.
  • Duplicated litellm.RateLimitError: prefixes are stripped.
  • A short plain-text detail is preserved (unchanged behaviour, guarded on purpose).
  • An unparsable payload is collapsed to one truncated line.
  • agenerate() condenses identically.

Full suite: 1558 passed. pre-commit run passes.

Additional Notes

  • Until retries are bounded, this message is raised on every attempt but never surfaces, because RateLimitError is retryable and the retry never ends. My PR for Issues with Gemini quickstart flow (retry behavior, model errors, and tool usage #266 fixes that; this PR is what the user finally sees. Either can merge first.
  • Both PRs touch module_llm.py at the same anchor points, so whichever merges second needs a trivial rebase. I will rebase immediately on request.
  • The message format is asserted by existing tests, so the wording is deliberately close to the current one.

_build_rate_limit_error() interpolates the provider detail verbatim, so
a quota error still reaches the modeller as the provider's raw JSON
report. For a Gemini free-tier 429 the rewritten message is 1308
characters over 36 lines, longer than the original error, while the
retryDelay the provider supplies is never shown.

Extract the human-readable message and the retry delay from the payload
and drop the rest, so the error reads as one line, and use the reported
delay in the guidance ("Please wait 27s") instead of the generic "a few
minutes". Payloads without a message, or that do not parse, are
collapsed to a single line and truncated, which leaves short plain-text
details unchanged. Repeated "litellm.RateLimitError:" prefixes are
stripped rather than the first one only.

Gemini free-tier 429: 1308 chars / 36 lines -> 317 chars / 1 line.

Fixes mesa#257
@coderabbitai

coderabbitai Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: a3a7b2ce-3c60-40a3-accd-63b2570d6145

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Poor error message when API rate limit is exceeded

1 participant