Skip to content

[macOS][Codex Desktop] Sudden ~60% increase in 5h quota usage overnight; large prompt-cache miss on small turn #51604

Description

@mehoney

What version of the Codex App are you using (From “About Codex” dialog)?

26.930.61225

What subscription do you have?

ChatGPT Plus

What platform is your computer?

Darwin 25.5.0 arm64 arm

What issue are you seeing?

Starting on October 7, 2026, Codex Desktop began consuming my 5-hour and weekly usage limits substantially faster than on October 6, despite no intentional change to my normal coding workflow.

I noticed this immediately because only a few relatively small coding tasks consumed about 30% of my 5-hour allowance and about 5% of the weekly allowance.

On October 6 I was routinely using GPT-6.1 Sol Medium for much larger tasks without this level of quota consumption.

I initially suspected reasoning effort and switched from GPT-6.1 Sol Medium to Low, but local Codex telemetry suggests that reasoning effort alone does not explain the change.

Consecutive-day comparison

October 6, 22:45–23:38

  • Total tokens: 9,909,509
  • Input tokens: 9,888,444
  • Cached input: 9,692,800
  • Uncached input: 195,644
  • Output tokens: 21,065
  • Reasoning output: 5,045
  • Cache rate: 98.02%
  • Token-count updates: 79
  • 5h usage: 23% -> 41% (+18 percentage points)
  • Weekly usage: 4% -> 6%

Model mix:

  • GPT-6.1 Sol Medium: 5,646,665 tokens
  • GPT-6.1 Sol Low: 4,262,844 tokens

October 7, 07:41–08:49

  • Total tokens: 12,661,074
  • Input tokens: 12,621,024
  • Cached input: 12,223,616
  • Uncached input: 397,408
  • Output tokens: 40,050
  • Reasoning output: 9,900
  • Cache rate: 96.85%
  • Token-count updates: 149
  • 5h usage: 0% -> 29% (+29 percentage points)
  • Weekly usage: 0% -> 5%

Model mix:

  • GPT-6.1 Sol Medium: 8,349,478 tokens
  • GPT-6.1 Sol Low: 4,311,596 tokens

Total token volume increased by about 28%, while 5-hour quota consumption increased by about 61%.

More importantly:

  • uncached input increased ~103%
  • output tokens increased ~90%

Suspicious prompt-cache behavior

One small turn on October 7 was particularly abnormal:

  • User prompt length: 525 characters
  • GPT-6.1 Sol Medium
  • 4 token updates
  • 3 exec calls + 1 request_user_input_async
  • Total tool-result size: 14,434 characters
  • Cached input: 435,968 tokens
  • Uncached input: 119,630 tokens
  • Output: 2,680 tokens
  • Reasoning output: 1,431 tokens

The ~119,630 uncached input tokens do not appear to be explained by the small user prompt or the ~14k characters of tool output.

This looks consistent with an unexpected prompt-cache miss/invalidation.

I also checked for:

  • auto-compaction: none
  • context/token-limit compaction: none
  • memory generation / gpt-5.6-terra: none
  • codex-auto-review: none
  • parent/subagent sessions: none detected

Possibly related to #48986 and #51088, but this reproduction is on Codex Desktop for macOS with GPT-6.1 Sol.

What steps can reproduce the bug?

This does not appear to be code-specific, so I do not have a minimal source-code snippet that reliably triggers it.

  1. Open an existing repository in Codex Desktop on macOS.
  2. Use GPT-6.1 Sol Medium or Low.
  3. Perform normal coding tasks involving file inspection and exec/JS tool calls.
  4. Observe the 5-hour usage percentage and local token telemetry.
  5. Compare quota consumption with previous sessions in the same repository and with the same workflow.

The regression became noticeable on October 7, 2026.

It is not consistently reproducible with one specific prompt. The observed pattern is:

  • substantially higher uncached input,
  • more token-count updates for similar work,
  • and at least one large apparent prompt-cache miss on a small turn.

I can provide additional sanitized telemetry if needed, but I am not posting raw JSONL/Desktop logs or session IDs because the sessions contain proprietary repository content.

What is the expected behavior?

With the same repository, workflow, model family, and similar coding tasks, quota consumption should remain roughly consistent. Existing context that is eligible for prompt caching should not unexpectedly become ~120k tokens of uncached input on a small turn without an identifiable cause.

Additional information

Starting October 7, the 5-hour quota began draining substantially faster. Uncached input approximately doubled, output nearly doubled, and at least one small turn produced ~119,630 uncached input tokens despite only a 525-character user prompt and ~14k characters of tool output.

Activity

  1. added
    appIssues related to the Codex desktop app
    on Oct 7, 2026
  2. added
    bugSomething isn't working
    rate-limitsIssues related to rate limits, quotas, and token usage reporting
    on Oct 7, 2026
  3. github-actions commented on Oct 7, 2026

    @github-actions
    Contributor
  4. mehoney commented on Oct 7, 2026

    @mehoney
    Author

    I reviewed #51088 and #50680.

    They appear related, but I am keeping this issue open because this reproduction is on Codex Desktop for macOS with GPT-6.1 Sol (Medium and Low), and I have a specific telemetry example suggesting a large prompt-cache miss:

    • 525-character user prompt
    • ~14k characters of total tool output
    • 435,968 cached input tokens
    • 119,630 uncached input tokens

    I also have a consecutive-day comparison showing the regression beginning on October 7.

    If maintainers confirm this has the same root cause as one of the linked issues, I am happy for this issue to be merged/closed as a duplicate.

  5. meowpizzacat commented on Oct 7, 2026

    @meowpizzacat

    I inspected a bounded sample of existing local Codex Desktop telemetry on my Mac and found a partially similar cache pattern, but I cannot corroborate the five-hour quota increase.

    Environment at the audit:

    • macOS; Codex Desktop 26.930.61225, build 13232
    • Bundled codex-cli 0.160.1
    • ChatGPT Pro
    • Selected workloads used gpt-6-astra with xhigh or max reasoning, rather than the GPT-6.1 Sol model in this report

    The closest comparison used the same workspace, gpt-6-astra with xhigh reasoning, and approximately 42 active minutes on each side. The sampled windows were October 6, 2026, 15:49–16:49 and October 7, 2026, 07:08–08:08, both America/Los_Angeles (PDT). Active minutes mean elapsed time within running turns, including tool waits.

    Metric October 6 baseline October 7
    Input tokens 15,361,419 11,373,906
    Cached input 15,083,904 11,028,736
    Uncached input 277,515 345,170
    Weighted cache rate 98.19% 96.97%
    Output tokens, including reasoning 58,123 41,162
    Completed user turns 1 1

    Uncached input increased 24.4%, despite lower total input and output. Baseline service-tier evidence was unavailable, so speed matching is not established.

    The strongest outlier occurred on October 7 at 07:26:43 PDT:

    • Previous response: 166,656 input; 165,888 cached; 768 uncached.
    • Outlier: 167,487 input; 22,912 cached; 144,575 uncached.
    • Cache rate fell from 99.54% to 13.68% across approximately 84 seconds.
    • The same turn, model and reasoning effort continued, without an intervening recorded compaction or instruction change.
    • The next response recovered to approximately 97.42% cached.

    Two images and 156 text characters arrived in the intervening tool result. Therefore, the small change in total input-token count does not establish that the cacheable input was unchanged. I cannot determine the server-side cause.

    The broader sample was mixed: another matched window had nearly unchanged cache rates (98.57% → 98.50%), while a separate workspace improved (93.94% → 96.24%).

    Quota attribution was not possible. The inspected Pro snapshots recorded only a 10,080-minute weekly window; no 300-minute window was present. The baseline included a saturated weekly counter and a subsequent reset-period change. Concurrent chats and automatic-review activity also overlapped.

    Accounting used individual response-usage records reconciled with cumulative increments. Separately recorded compaction usage was retained; no-change notices were distinguished; reasoning tokens were not added twice. The audit used existing evidence through October 7, 08:24:17 PDT and excluded its own usage.

    My evidence supports an isolated cache anomaly resembling part of this report, but does not establish a sustained regression, increased five-hour quota consumption, or the same underlying bug.

  6. darkautism commented on Oct 9, 2026

    @darkautism

    I have cache miss too. It only 33% in cache in a session.
    Only sol has this issue. Luna cache worked.

  7. jordanovt commented on Oct 10, 2026

    @jordanovt

    I am seeing a very similar issue.

    Over the last two days, my Codex / ChatGPT usage allowance has started draining much faster than before, while my workflow has remained broadly comparable.

    I am not looking for generic advice about reducing token usage. The question is more specific:

    • did anything change in quota accounting, model routing, reasoning-token accounting, prompt-cache behavior, tool/sub-agent execution, retries, or internal multipliers?
    • is allowance consumption based only on user-visible input/output, or can hidden OpenAI-controlled mechanisms materially affect the effective burn rate?
    • if hidden reasoning, background work, retries, tool calls, or sub-agent activity consume quota, how can users inspect or limit that consumption?

    This is important from a trust and predictability perspective. Paying users need to understand whether the allowance is being consumed by actions they explicitly control, or by internal behavior that is not visible in the normal UI.

    The pattern reported in this issue is consistent with what I am seeing: sudden materially higher quota burn over a very short period, without a matching intentional change in workflow.

  8. raviknits commented on Oct 11, 2026

    @raviknits

    If it helps with the before/after comparison: this offline script reads ~/.codex/sessions and prints what your recent sessions sent to the model and how much of it was sent again, so you don't have to go through rollout files by hand:

    curl -fsSL https://densilo.com/tokens.py | python3 - --days 7
    

    It can't see server-side cache hits, but it does show whether the client started sending more. (Disclosure: I build Densilo, and the script uploads nothing.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appbugSomething isn't workingrate-limitsIssues related to rate limits, quotas, and token usage reporting

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions