Skip to content

feat(codex): report free rate-limit reset credits - #430

Merged
MagicalTux merged 3 commits into
KarpelesLab:masterfrom
rikbrown:feat/codex-reset-credit-count
Sep 19, 2026
Merged

MagicalTux merged 3 commits into
KarpelesLab:masterfrom
rikbrown:feat/codex-reset-credit-count

Conversation

@rikbrown

@rikbrown rikbrown commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

The gap

OpenAI grants ChatGPT/Codex subscriptions occasional free rate-limit reset credits. Redeeming one clears that account's spent rate-limit windows before their reset. The Codex client exposes this as a manual action in its usage screen.

TeamClaude did not report these credits. A pooled account could sit out a spent weekly window while an expiring credit went unused.

What changed

Every existing quota surface now shows the credit count. Nothing redeems a credit.

No extra request. The /wham/usage payload that the Codex quota probe already reads includes top-level rate_limit_reset_credits. src/codex-usage.js parses it into quota.resetCredits.

Two separate counts. available_count is the count an account holds and is shown on every surface. applicable_available_count is upstream's count that would reset a window now; it is 0 when no window is eligible. Neither says whether a particular credit can actually be spent, which only that account's credit rows state. Missing or malformed counts are dropped rather than treated as zero.

Persistence. resetCredits joins PERSISTED_QUOTA_FIELDS because the usage probe is off by default. A payload without credits retains the previous reading, which is stamped so its age is visible.

Display. The TUI account row shows RC1 (ASCII, as with the spend tag, because the row is cell-budgeted), teamclaude status gains a Reset line that says when none applies yet, and dashboard cards get a badge.

The wire format

rate_limit_reset_credits is not part of the public OpenAI API and is documented nowhere. It was recovered from the Codex binary and verified against a live account:

"rate_limit_reset_credits": { "available_count": 1, "applicable_available_count": 0 }

Verification

  • npm test — 1874 tests pass, including 6 new tests in test/codex-reset-credits.test.js and 3 in test/codex-usage.test.js.
  • npm run lint, npm run typecheck — clean.
  • node scripts/typecheck-strict.mjs --base upstream/master — 2 below the base. A CodexLearnedQuota typedef for Codex-only fields written by applyCodexUsageData also fixes two pre-existing diagnostics.

No test reaches the network.

What this does not do

Spend a credit. That irreversible action is a separate stacked change, with its own switch, default off.

🤖 Generated with Claude Code

OpenAI occasionally grants a ChatGPT subscription a free rate-limit reset
credit: redeeming one clears that account's spent rate-limit windows ahead of
their own reset. Codex's own client spends one by hand, from its usage screen —
but nothing here said an account even held one, so a pooled account could sit
out the rest of its week with an unspent credit in hand. Credits expire.

The count comes free. `/wham/usage` — the payload the quota probe already reads
— carries `rate_limit_reset_credits` at the top level, so reporting it costs no
request of its own: it is parsed beside the rest of that payload and lands on
the account as `quota.resetCredits`.

Two counts, kept apart on purpose. `available_count` is what the account HOLDS,
and is the number every surface draws. `applicable_available_count` is
upstream's own view of how many would reset a window right now, which is 0
whenever no window is currently eligible. Neither is a verdict on whether a
given credit could actually be spent — that is stated only by the account's own
credit rows — so the two are never conflated into one "you can redeem this"
number that nothing ever claimed.

A missing or malformed count is dropped rather than read as zero, the same way
a zeroed window already is: "none" and "we were not told" have different
consequences, and only one of them is a fact.

`resetCredits` joins PERSISTED_QUOTA_FIELDS because the usage probe is off by
default — without that, a restart forgets the one thing that says a credit
exists at all until something next happens to read the usage endpoint. For the
same reason a payload that mentions no credits leaves the last reading alone
rather than blanking it, and the reading is stamped, so how much the number is
still worth can be told from its age.

An operator sees it as `RC1` on the TUI account row (ASCII, for the same reason
the spend tag is: the row is budgeted to the cell), a `Reset` line in
`teamclaude status` that says when none would apply yet, and a badge on the
dashboard card.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@rikbrown rikbrown changed the title feat(codex): report the free rate-limit reset credits an account holds feat(codex): report free rate-limit reset credits Sep 19, 2026
MagicalTux and others added 2 commits September 20, 2026 08:35
The credit counts from /wham/usage were drawn verbatim as `RC<n>` on a
width-budgeted TUI row; they are now truncated and capped at 99, and a
non-finite or negative value is dropped. `seenAt` was stamped but never read,
so with the probe off a redeemed or expired credit stayed on screen forever:
the status line now says how old the reading is, and the status screen, TUI
row and dashboard badge all hide a reading more than 7 days old.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@MagicalTux
MagicalTux merged commit 07fbb13 into KarpelesLab:master Sep 19, 2026
5 checks passed
@MagicalTux MagicalTux mentioned this pull request Sep 20, 2026
MagicalTux added a commit that referenced this pull request Sep 20, 2026
Thirty-six commits since 1.1.20. Several change what a client sees on a
failure, so read the first section before upgrading a shared deployment.

Behaviour changes
  #439 a 401 from upstream is no longer relayed to the client: the request
       fails over like a 403, and an API-key account or an OAuth account with
       no refresh token leaves rotation (it used to be picked again on every
       request). With nothing left the client gets the proxy's own error
  #439 a path under /teamclaude/ that no control route claims — a typo, or the
       wrong verb — answers 404 locally instead of being forwarded upstream
       under a fleet credential
  #429 the synthetic 429's retry-after is the real reset of the windows
       blocking the request's candidate accounts, not a flat 60s, and the
       message counts only those candidates; #408 names accounts that need a
       re-login instead of calling them "at quota"
  #438 session pins are per conversation (session id plus a digest of the
       first message), so a session's subagents spread across accounts.
       `sessions.items[].id` in status is the composite key, load is counted
       per conversation, and a persisted concurrency cap re-learns
  #378 a `thread: continue` bound for a per-account third-party upstream is
       refused with the 400 Anthropic gives, so the client resends the whole
       conversation; `messageThreads: true` opts a relay out
  #434 a 429 whose x-codex-* headers show a spent account-wide window holds
       the account like an Anthropic rejection; a spent model-scoped bucket
       only moves the request
  #437 a Codex response head is awaited for five minutes (Anthropic unchanged)
  #411 idle keep-alive connections are held 120s on both listeners
  #389 with session distribution on, requests carrying no session id rotate on
       a cursor of their own instead of all resting on the current account
  #405 `defaultClientMode` ("mitm" | "base-url") sets what `run` and `env` do
       without a flag; `--mitm` / `--no-mitm` decide per launch, and in
       base-URL mode `env` unsets a stale proxy export naming this proxy
  #439 route `--bucket` is validated; an array `switchThreshold` reads as the
       default with one line saying so

Features
  #419 an MCP management endpoint at POST /teamclaude/mcp, off unless
       `proxy.mcp` is "read" or "full"; a named client key is read-only even
       in full mode, and with no proxy key it serves only this machine
  #428 per-account `switchThreshold`, a number or a per-bucket table
  #406 a Claude+Codex pool is drawn as two panes on a wide terminal, each with
       its own current marker; #392 names the provider in a mixed list; #418
       lets the operator arrange the list (`displayOrder`); #376 draws
       loopback-served accounts last; #435 shows the percentage beside a bar's
       countdown; #394 shows the running version in the header
  #430 free Codex rate-limit reset credits in status, the TUI and the dashboard
  #385 `proxy.terminalOnly` tunnels chatgpt.com so ChatGPT Desktop stays out

Fixes
  #404 a TUI paint can no longer block the proxy (stdout non-blocking, frames
       dropped while the terminal is behind); #410 a dead terminal no longer
       takes the proxy with it, and SIGHUP shuts down cleanly
  #433 token usage is booked from Codex Responses streams
  #386 #387 #388 the Codex five-hour window is read from the model-scoped
       family and the usage probe, and extra limits are named from their entries
  #431 a headerless 429 that follows the request is retried once
  #432 #439 startup and collaborator log lines reach the TUI's activity pane
       and log file instead of the covered terminal
  #415 #403 #439 reload mirrors `priority`, `disabled`, `stripRequestFields`
       onto the config entry, and a reload during a removal does not re-add it
  #439 the Host check uses the address actually bound; sx.org calls time out
  #381 the dashboard polls status before asking for a key
  #403 session outcome accounting classifies the decoded path; account names in
       daemon log lines are sanitised

Tooling
  #371 #372 #373 `npm run typecheck` (tsc over the JS sources) in CI, with a
       strict-mode ratchet: per-file strict diagnostics may not grow past the
       pre-merge commit (2006 at introduction, 1735 now)
  #401 docker workflow actions bumped

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
rikbrown added a commit to rikbrown/teamclaude that referenced this pull request Sep 20, 2026
Builds on the reading of `quota.resetCredits` (KarpelesLab#430): this is the one action
that count makes possible, and it is off unless an operator arms it.

A Codex account whose weekly window has run out is often holding the credit
that would clear it. Codex's own client spends one by hand, from its usage
screen, so a pooled account stayed walled for the rest of its week while the
credit sat unspent and expiring.

The trigger is the refusal itself. Once every Codex account reads over
threshold, selection turns the request away BEFORE choosing one: the proxy
answers "(none available)" and no upstream request is made — so there is no 429
to hang this off, and on a spent pool that is what happens to every request.
The hook therefore sits in forwardRequest's `!account` branch, ahead of the
retry-after measurement and the hold/retry paths, and is offered only the
accounts whose `unavailableReason` a reset would actually clear: `quota` and
`throttled`, never `disabled`, `capped`, `entitlement`, `error` or `route`, all
of which survive a cleared window and would take a credit for nothing. A
decline or a throw leaves the refusal byte-for-byte what it was; a success
re-selects, costing no retry from the budget because nothing was ever sent.

The two routes are not part of the public API and are documented nowhere;
these were recovered from the Codex binary and then verified live:

  GET  /wham/rate-limit-reset-credits           detail rows: id, status,
       expires_at, is_supported_by_plan. ISO-8601 STRINGS here, while the Rust
       app-server layer states the same fields as epoch seconds — the two must
       not be parsed the same way
  POST /wham/rate-limit-reset-credits/consume   {redeem_request_id, credit_id?}

Two traps in that last one, both covered by tests: the verdict field is `code`,
not the `outcome` the app-server layer uses, and it answers 200 when it refuses
as readily as when it works. Classify on the body, never the status — reading
"you hold no credits" as a redemption that worked is not recoverable.

Spending a credit is irreversible and they are scarce, so the policy is
deliberately mean. A redemption needs the WEEKLY window spent — never the 5-hour
one, which comes back on its own within hours while the weekly one walls an
account off for days — and then either the whole Codex pool is dry, so the
credit actually unblocks work rather than topping up an account rotation would
have stepped past, or the credit expires within three days and holding it costs
more than spending it. A credit the plan cannot spend is filtered before any
expiry reasoning, so it can never be what makes an expiring-credit decision look
justified.

At most one credit per dry pool, which takes three guards rather than one.
Refusals arriving together join a single attempt. An attempt that spent a
credit — or that failed in a way that cannot rule out having spent one — holds
the WHOLE fleet off, not only the account it touched: an attempt walks the
pool, so a sibling that never touched the endpoint is just as able to spend the
second credit, and the alternative (the redeemed account reading available
again) leans on a quota re-read that is allowed to fail. And the pool is
re-resolved per decision, so once that re-read does land, the account it
returned to service is what makes every sibling answer "another Codex account
can still serve". `ctx.resetRedeemTried` bounds it to one attempt per request,
so a redeem that reports success but leaves the account unselectable — upstream
not yet caught up with its own reset — costs that request one re-selection
rather than a loop.

`redeem_request_id` is an idempotency key. A consume whose verdict never
arrived may still have been acted on, so it replays its key instead of minting
a fresh one, and `already_redeemed` is upstream answering for the POST that did
land — which is why it counts as a success.

Off unless armed: `autoRedeemResets` is a fleet switch, default false, in
`createDefaultConfig` and `config.example.json`, toggleable from the TUI
settings screen and applied live on reload so it can be killed without a
restart. The per-account `accounts[].autoRedeemReset` is a veto only — it can
say never for this account, never yes.

The attempt is awaited inline, before the re-selection, because a dry pool has
nowhere to rotate to: the redemption IS the recovery for that request. It is
bounded by a single 10s deadline across token refresh, detail fetch, consume
and the re-read together, leaving the rest of the client's response-head
budget for the retry it exists to enable. The refresh is raced against that
deadline rather than measured after it returns: `ensureTokenFresh` takes no
timeout of its own and every concurrent refusal joins the same one, so a check
that runs afterwards can only report a promise already broken.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MagicalTux added a commit that referenced this pull request Sep 25, 2026
…ry (#436)

Builds on the reading of `quota.resetCredits` (#430): this is the one action
that count makes possible, and it is off unless an operator arms it.

A Codex account whose weekly window has run out is often holding the credit
that would clear it. Codex's own client spends one by hand, from its usage
screen, so a pooled account stayed walled for the rest of its week while the
credit sat unspent and expiring.

The trigger is the refusal itself. Once every Codex account reads over
threshold, selection turns the request away BEFORE choosing one: the proxy
answers "(none available)" and no upstream request is made — so there is no 429
to hang this off, and on a spent pool that is what happens to every request.
The hook therefore sits in forwardRequest's `!account` branch, ahead of the
retry-after measurement and the hold/retry paths, and is offered only the
accounts whose `unavailableReason` a reset would actually clear: `quota` and
`throttled`, never `disabled`, `capped`, `entitlement`, `error` or `route`, all
of which survive a cleared window and would take a credit for nothing. A
decline or a throw leaves the refusal byte-for-byte what it was; a success
re-selects, costing no retry from the budget because nothing was ever sent.

The two routes are not part of the public API and are documented nowhere;
these were recovered from the Codex binary and then verified live:

  GET  /wham/rate-limit-reset-credits           detail rows: id, status,
       expires_at, is_supported_by_plan. ISO-8601 STRINGS here, while the Rust
       app-server layer states the same fields as epoch seconds — the two must
       not be parsed the same way
  POST /wham/rate-limit-reset-credits/consume   {redeem_request_id, credit_id?}

Two traps in that last one, both covered by tests: the verdict field is `code`,
not the `outcome` the app-server layer uses, and it answers 200 when it refuses
as readily as when it works. Classify on the body, never the status — reading
"you hold no credits" as a redemption that worked is not recoverable.

Spending a credit is irreversible and they are scarce, so the policy is
deliberately mean. A redemption needs the WEEKLY window spent — never the 5-hour
one, which comes back on its own within hours while the weekly one walls an
account off for days — and then either the whole Codex pool is dry, so the
credit actually unblocks work rather than topping up an account rotation would
have stepped past, or the credit expires within three days and holding it costs
more than spending it. A credit the plan cannot spend is filtered before any
expiry reasoning, so it can never be what makes an expiring-credit decision look
justified.

At most one credit per dry pool, which takes three guards rather than one.
Refusals arriving together join a single attempt. An attempt that spent a
credit — or that failed in a way that cannot rule out having spent one — holds
the WHOLE fleet off, not only the account it touched: an attempt walks the
pool, so a sibling that never touched the endpoint is just as able to spend the
second credit, and the alternative (the redeemed account reading available
again) leans on a quota re-read that is allowed to fail. And the pool is
re-resolved per decision, so once that re-read does land, the account it
returned to service is what makes every sibling answer "another Codex account
can still serve". `ctx.resetRedeemTried` bounds it to one attempt per request,
so a redeem that reports success but leaves the account unselectable — upstream
not yet caught up with its own reset — costs that request one re-selection
rather than a loop.

`redeem_request_id` is an idempotency key. A consume whose verdict never
arrived may still have been acted on, so it replays its key instead of minting
a fresh one, and `already_redeemed` is upstream answering for the POST that did
land — which is why it counts as a success.

Off unless armed: `autoRedeemResets` is a fleet switch, default false, in
`createDefaultConfig` and `config.example.json`, toggleable from the TUI
settings screen and applied live on reload so it can be killed without a
restart. The per-account `accounts[].autoRedeemReset` is a veto only — it can
say never for this account, never yes.

The attempt is awaited inline, before the re-selection, because a dry pool has
nowhere to rotate to: the redemption IS the recovery for that request. It is
bounded by a single 10s deadline across token refresh, detail fetch, consume
and the re-read together, leaving the rest of the client's response-head
budget for the retry it exists to enable. The refresh is raced against that
deadline rather than measured after it returns: `ensureTokenFresh` takes no
timeout of its own and every concurrent refusal joins the same one, so a check
that runs afterwards can only report a promise already broken.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mark Karpeles <magicaltux@gmail.com>
rikbrown added a commit to rikbrown/teamclaude that referenced this pull request Oct 7, 2026
…redits

Upstream KarpelesLab#430 and KarpelesLab#436 now report and redeem free rate-limit reset credits,
so what is left of this change is the fork's side of it.

Persists quota.planType and quota.codexModelBuckets, which were learned and
then dropped on every restart, and clamps a restored bucket table to the 32
newest — the writers cap, but the eviction only runs when a new slug
arrives, so an over-long table would otherwise stand.

docs/openai.md gains the reset-credit section, for a pool reached through
the Codex sidecar.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants