Skip to content

A 401 is relayed to the client and the account stays in rotation: the only 401 branch is gated on account.refreshToken, and unlike 403 it never fails over #412

Description

@jeffheartmp

Summary

403 rotates the account out. 5xx hops once. 429 throttles and switches. 401 does neither — its only branch is a re-auth that is gated on account.refreshToken, so on a pool whose OAuth accounts carry no stored refresh token the 401 has no handler at all: it falls through to the relay, the account stays active, and it is picked again on the very next request.

Measured on 1.1.20: one account returned 4,124 consecutive 401s over ~20 hours, every one relayed to the client, while three healthy accounts in the same pool served none of them.

Where

src/server.js — the 403 branch rotates:

if (upstreamRes.status === 403 && !res.headersSent) {
  ...
  ctx.tried.add(account.index);
  console.error(`[TeamClaude] 403 on "${account.name}"; upstream refused the account credential${cooldown}`);
  return forwardRequest(req, res, body, accountManager, upstream, retryCount + 1, ...);
}

The 401 branch does not — and it is gated:

if (upstreamRes.status === 401 && account.type === 'oauth' && account.refreshToken
    && retryCount < maxRetries && !ctx.reauthed.has(account.index)) {
  ctx.reauthed.add(account.index);
  ...
  await accountManager.ensureTokenFresh(account.index, true);
  return forwardRequest(...);
}
// falls through to relay — no ctx.tried.add, no failover

If account.refreshToken is falsy the whole block is skipped. There is no ctx.tried.add(account.index) + forwardRequest for 401 the way there is for 403, so the request is answered with the 401 and the account is left in rotation.

Why the existing comment does not cover it

The comment above that branch says:

If the refresh is itself rejected the refresh token is dead too: ensureTokenFresh marks the account errored, and the retry's status check rotates to another account.

That is true only when a refresh actually runs. The rotation is a side effect of ensureTokenFresh's _deadRefreshToken guard setting status = 'error'. With no stored refresh token there is no refresh, no guard, no error status, and therefore no rotation — the one path that takes a bad-credential account out of service never executes.

This matters for any deployment where access tokens are provisioned externally and pushed in (updateAccountTokens / teamclaude import) rather than refreshed by the proxy from a stored refresh token. That is a supported shape — Updated tokens for account "…" fires normally on this pool — but it silently disables the only 401 handling there is.

Evidence

Accounts as configured (all OAuth, none with a stored refresh token):

{"name":"account-a@example.com","type":"oauth","hasRefresh":false,"hasApiKey":false,"upstream":null}
{"name":"account-b@example.com","type":"oauth","hasRefresh":false,"hasApiKey":false,"upstream":null}
{"name":"account-c@example.com","type":"oauth","hasRefresh":false,"hasApiKey":false,"upstream":null}

One client's traffic over 11 days, by account and status:

account-b@example.com   401   4124
account-b@example.com   200    755
account-a@example.com   200    323
account-c@example.com   200     54
(none available)        429    320

All 4,124 401s are a single unbroken run (longest-consecutive-401 run == total 401 count), start to finish on the same account. Per-session follow-up behaviour:

401 followed by a retry on the SAME account:       3853
401 followed by a retry on a DIFFERENT account:       0

The re-auth path never ran, and no account was ever marked error:

grep -c 'forcing refresh and retrying'      -> 0
grep -c 'rejected refresh token'            -> 0
grep -c 'invalid_grant'                     -> 0
grep -c 'upstream refused the account'      -> 5     (the 403 path, working)
grep -c 'Quota rejection'                   -> 408   (the 429 path, working)

Throughout the whole ~20h window /teamclaude/status reported the failing account as:

{"name":"account-b@example.com","status":"active","unavailable":null,"quota":{"unified5h":0.02,"unified7d":0.12}}

active, nothing unavailable, and near-zero quota use — precisely because its requests were failing instantly and burning nothing. Every operator-facing signal said the account was healthy.

Impact

One account whose credential upstream has stopped accepting takes down whichever clients the cursor happens to land on, for as long as the credential stays bad, with no log line and no status change to point at it. Clients that land on a different account are unaffected, which makes it read as an intermittent client-side fault rather than one account being dead.

Expected

  1. Give 401 the 403 treatment: ctx.tried.add(account.index) and return forwardRequest(...) so the request is retried on another account instead of being relayed.
  2. Do not gate rotation on account.refreshToken. Keep the forced refresh gated on it — that part is correct — but the failover should happen whether or not a refresh was possible.
  3. After N consecutive 401s on an account, mark it unavailable with a cooldown (the markEntitlementDenied shape already in the 403 path) so /teamclaude/status and the TUI stop reporting it as active, and so the exhausted-message can name it.

Point 3 is the half of #407 that applies here: that issue is about the synthetic 429 wording when an account is already status=error. This one is about an account that never reaches error in the first place.

Environment

  • @karpeleslab/teamclaude 1.1.20 (latest at time of filing), installed globally, systemd unit per pool
  • 3 OAuth accounts, tokens provisioned externally, no stored refresh tokens
  • Linux, Node 22

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions