Repository navigation
[macOS][Codex Desktop] Sudden ~60% increase in 5h quota usage overnight; large prompt-cache miss on small turn #51604
Description
Activity
- addedappIssues related to the Codex desktop appIssues related to the Codex desktop app
on Oct 7, 2026 - addedbugSomething isn't workingSomething isn't workingrate-limitsIssues related to rate limits, quotas, and token usage reportingIssues related to rate limits, quotas, and token usage reporting
on Oct 7, 2026 Potential duplicates detected. Please review them and close your issue if it is a duplicate.
- Codex quota drains abnormally fast even on trivial cached prompts #51088
- GPT-5.6 Light consuming ~20% of 5-hour allowance on minimal read-only task #50680
Powered by Codex Action
They appear related, but I am keeping this issue open because this reproduction is on Codex Desktop for macOS with GPT-6.1 Sol (Medium and Low), and I have a specific telemetry example suggesting a large prompt-cache miss:
- 525-character user prompt
- ~14k characters of total tool output
- 435,968 cached input tokens
- 119,630 uncached input tokens
I also have a consecutive-day comparison showing the regression beginning on October 7.
If maintainers confirm this has the same root cause as one of the linked issues, I am happy for this issue to be merged/closed as a duplicate.
meowpizzacat commented
on Oct 7, 2026 More actionsI inspected a bounded sample of existing local Codex Desktop telemetry on my Mac and found a partially similar cache pattern, but I cannot corroborate the five-hour quota increase.
Environment at the audit:
- macOS; Codex Desktop 26.930.61225, build 13232
- Bundled codex-cli 0.160.1
- ChatGPT Pro
- Selected workloads used gpt-6-astra with xhigh or max reasoning, rather than the GPT-6.1 Sol model in this report
The closest comparison used the same workspace, gpt-6-astra with xhigh reasoning, and approximately 42 active minutes on each side. The sampled windows were October 6, 2026, 15:49–16:49 and October 7, 2026, 07:08–08:08, both America/Los_Angeles (PDT). Active minutes mean elapsed time within running turns, including tool waits.
Metric October 6 baseline October 7 Input tokens 15,361,419 11,373,906 Cached input 15,083,904 11,028,736 Uncached input 277,515 345,170 Weighted cache rate 98.19% 96.97% Output tokens, including reasoning 58,123 41,162 Completed user turns 1 1 Uncached input increased 24.4%, despite lower total input and output. Baseline service-tier evidence was unavailable, so speed matching is not established.
The strongest outlier occurred on October 7 at 07:26:43 PDT:
- Previous response: 166,656 input; 165,888 cached; 768 uncached.
- Outlier: 167,487 input; 22,912 cached; 144,575 uncached.
- Cache rate fell from 99.54% to 13.68% across approximately 84 seconds.
- The same turn, model and reasoning effort continued, without an intervening recorded compaction or instruction change.
- The next response recovered to approximately 97.42% cached.
Two images and 156 text characters arrived in the intervening tool result. Therefore, the small change in total input-token count does not establish that the cacheable input was unchanged. I cannot determine the server-side cause.
The broader sample was mixed: another matched window had nearly unchanged cache rates (98.57% → 98.50%), while a separate workspace improved (93.94% → 96.24%).
Quota attribution was not possible. The inspected Pro snapshots recorded only a 10,080-minute weekly window; no 300-minute window was present. The baseline included a saturated weekly counter and a subsequent reset-period change. Concurrent chats and automatic-review activity also overlapped.
Accounting used individual response-usage records reconciled with cumulative increments. Separately recorded compaction usage was retained; no-change notices were distinguished; reasoning tokens were not added twice. The audit used existing evidence through October 7, 08:24:17 PDT and excluded its own usage.
My evidence supports an isolated cache anomaly resembling part of this report, but does not establish a sustained regression, increased five-hour quota consumption, or the same underlying bug.
I have cache miss too. It only 33% in cache in a session.
Only sol has this issue. Luna cache worked.jordanovt commented
on Oct 10, 2026 More actionsI am seeing a very similar issue.
Over the last two days, my Codex / ChatGPT usage allowance has started draining much faster than before, while my workflow has remained broadly comparable.
I am not looking for generic advice about reducing token usage. The question is more specific:
- did anything change in quota accounting, model routing, reasoning-token accounting, prompt-cache behavior, tool/sub-agent execution, retries, or internal multipliers?
- is allowance consumption based only on user-visible input/output, or can hidden OpenAI-controlled mechanisms materially affect the effective burn rate?
- if hidden reasoning, background work, retries, tool calls, or sub-agent activity consume quota, how can users inspect or limit that consumption?
This is important from a trust and predictability perspective. Paying users need to understand whether the allowance is being consumed by actions they explicitly control, or by internal behavior that is not visible in the normal UI.
The pattern reported in this issue is consistent with what I am seeing: sudden materially higher quota burn over a very short period, without a matching intentional change in workflow.
If it helps with the before/after comparison: this offline script reads ~/.codex/sessions and prints what your recent sessions sent to the model and how much of it was sent again, so you don't have to go through rollout files by hand:
curl -fsSL https://densilo.com/tokens.py | python3 - --days 7It can't see server-side cache hits, but it does show whether the client started sending more. (Disclosure: I build Densilo, and the script uploads nothing.)
What version of the Codex App are you using (From “About Codex” dialog)?
26.930.61225
What subscription do you have?
ChatGPT Plus
What platform is your computer?
Darwin 25.5.0 arm64 arm
What issue are you seeing?
Starting on October 7, 2026, Codex Desktop began consuming my 5-hour and weekly usage limits substantially faster than on October 6, despite no intentional change to my normal coding workflow.
I noticed this immediately because only a few relatively small coding tasks consumed about 30% of my 5-hour allowance and about 5% of the weekly allowance.
On October 6 I was routinely using GPT-6.1 Sol Medium for much larger tasks without this level of quota consumption.
I initially suspected reasoning effort and switched from GPT-6.1 Sol Medium to Low, but local Codex telemetry suggests that reasoning effort alone does not explain the change.
Consecutive-day comparison
October 6, 22:45–23:38
Model mix:
October 7, 07:41–08:49
Model mix:
Total token volume increased by about 28%, while 5-hour quota consumption increased by about 61%.
More importantly:
Suspicious prompt-cache behavior
One small turn on October 7 was particularly abnormal:
The ~119,630 uncached input tokens do not appear to be explained by the small user prompt or the ~14k characters of tool output.
This looks consistent with an unexpected prompt-cache miss/invalidation.
I also checked for:
Possibly related to #48986 and #51088, but this reproduction is on Codex Desktop for macOS with GPT-6.1 Sol.
What steps can reproduce the bug?
This does not appear to be code-specific, so I do not have a minimal source-code snippet that reliably triggers it.
The regression became noticeable on October 7, 2026.
It is not consistently reproducible with one specific prompt. The observed pattern is:
I can provide additional sanitized telemetry if needed, but I am not posting raw JSONL/Desktop logs or session IDs because the sessions contain proprietary repository content.
What is the expected behavior?
With the same repository, workflow, model family, and similar coding tasks, quota consumption should remain roughly consistent. Existing context that is eligible for prompt caching should not unexpectedly become ~120k tokens of uncached input on a small turn without an identifiable cause.
Additional information
Starting October 7, the 5-hour quota began draining substantially faster. Uncached input approximately doubled, output nearly doubled, and at least one small turn produced ~119,630 uncached input tokens despite only a 525-character user prompt and ~14k characters of tool output.