Repository navigation
feat(accounts): per-account egress proxy (http, socks4/4a, socks5/5h) - #441
Conversation
… and request forwarding
…xy is unreachable
…at account's routing
…hake, and IPv6 literal targets reach SOCKS as addresses
…nstead of adding the account unrouted
…s on it, with --no-check to skip and routing --check on demand
…hrough its stored routing
…is taller than the terminal
…ed on this branch
…ver and cooldown, the TUI row, and accounts[].routing in the config reference
…ort whose profile lookup went around the account's proxy
…stored routing instead of naming a proxy host "none"
…not a proxy host named none
…erid does not print as user:***
… judged before they are sized, mapped IPv6 tails routingAgent now gives the tunnel and the TLS handshake over it 20s each (TEAMCLAUDE_ROUTING_TIMEOUT_MS overrides it, for tests), where before they ran on the 30s CONNECT default. The OAuth refresh wraps its fetch in AbortSignal.timeout(30s), so against a black-holed proxy that signal won the race and the failure surfaced as a generic timeout, retried as transient on the same dead proxy, instead of the ROUTING_FAILED that fails the account over and arms its cooldown. A test with a proxy that accepts the connection and never answers pins the failover and the hold. SOCKS5: the reply's version and REP are checked before the reply is sized by ATYP, so a refusal that carries ATYP 0 (or stops after REP) reads as the refusal it is rather than "unknown address type" or a wait for bytes that never come. SOCKS4: the VN byte is checked before CD is trusted, so a non-SOCKS answer (an HTTP proxy's text) is named as such. ipv6Bytes now builds the 16 bytes from forward-target.js's v6Groups (exported for it), which folds an embedded IPv4 tail (::ffff:1.2.3.4, how a dual-stack resolver reports a mapped address) into two groups; the hand-rolled splitter packed the dotted quad as one garbage group. checkRouting parses the URL inside its try, so it resolves with the failure for an unparseable upstream too, as its contract says. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…listener The MITM listener intercepts api.anthropic.com, so a routing pointing at this server's own address would tunnel every request for the account straight back into forwardRequest. accountRouting() now takes the listener (upstream-proxy.js localListener / isSelfProxy, the same guard the fleet upstreamProxy has) and drops such a value with an "ignoring routing — that address is this server" line; the AccountManager carries it as an opts `listener` so accounts added at runtime get the same check, and the config reload passes the running config's. The write sites refuse it before the proxy test, which this server would pass: the routing command, login/import --routing, the MCP set_account_routing tool and the TUI prompt. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d a self-address routing is refused Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts: # docs/configuration.md # docs/usage.md # src/account-manager.js # src/index.js # src/sync-accounts.js
|
I reviewed this closely and it is going in. Two things were done on the branch, so you know what changed under you:
CI will run on the updated branch; I will merge once it is green. |
Thirty-two commits since 1.1.21. Two change routing on an existing config without an opt-in (#480, #481); the rest is opt-in, additive, or display. Behaviour changes #481 an API-key account's 401 is a cooldown, not a permanent `error`: 1 min, then 5, 15 and 60 for every further rejection with no success in between; any 2xx/3xx resets it. The request still fails over and the client never sees the 401. OAuth accounts are unchanged #480 with session distribution on, requests carrying no session id stay within the top priority tier, so a fallback gateway no longer answers Claude Code's bootstrap and connector calls #470 a 200 whose SSE stream reports a provider failure before any output (`server_is_overloaded`, `response.failed`) fails over once, like a status-shaped failure would #465 a reload removes running accounts whose config entry is gone from disk, so `teamclaude remove` from another shell takes effect at once #460 `import` refuses an account whose token upstream has definitively rejected (401/403), even with `--name`; a 5xx or timeout still imports Rename #483 the project is being renamed to TeamRouter (#72). This release accepts the new name everywhere the old one is read and changes nothing an install has on disk: `teamrouter` runs the same CLI, every `TEAMCLAUDE_*` variable is also read as `TEAMROUTER_*` (which wins when both are set), every `/teamclaude/…` control route also answers at `/teamrouter/…`, and `~/.config/teamrouter.json` is used when it exists Features #441 per-account egress proxy (`accounts[].routing`: http, socks4/4a, socks5/5h) for refresh, probes and requests; `login --routing`, `teamclaude routing set/show/clear`, a connection check before it is relied on, and a short hold when the proxy is unreachable #427 `accounts[].allowExtraUsage: true` lets a paid extra-usage account serve once every account is past its threshold, instead of a 429 #466 `accounts[].maxSpend`, a money cap judged against the month-to-date extra-usage spend upstream reports; the TUI shows what an account has billed #436 `autoRedeemResets` spends a free Codex rate-limit reset credit when the Codex pool runs dry (off by default) #482 `advisorEligibility: "strict" | "prefer"`; when the advisor model narrows selection to a subset of the fleet the log says so, and status carries the reading (`advisorNarrowing`) #478 `stripOverageHeaders` drops another org's per-organization billing headers from responses, for a pool spanning several orgs (#476) #471 `quota.unified5hSeenAt` / `unified7dSeenAt` in status: when upstream last stated each shared window #446 client and dimension usage for the last 5h and 24h in status and the dashboard, resumed across restarts #458 #459 #461 the dashboard sets the switch threshold, enables/disables and reprioritizes an account, and has a light theme remembered per browser #464 `l` in the TUI signs an account in `error` in again from the dashboard #457 status records which Codex limit meters each model (`quota.codexModelLimits`) #442 `quotaBarPercent` drops the percentage beside a TUI bar's countdown #451 `stripRequestFields` takes `content.<block type>` to drop content blocks a strict Anthropic-compatible upstream rejects #469 `proxy.mcp` schemas declare their item types, the write audit line records what happened, and the write queue has a depth (#447–#450) Fixes #477 a refused WebSocket handshake whose headers all drop is relayed as a well-formed head instead of a blank line and body bytes #474 two members of one ChatGPT workspace are told apart by user id, so a second `login --codex` no longer replaces the first #469 a Codex Responses stream with no Content-Type is relayed as a stream and booked; thread repair on the global upstream; TUI settings and status gaps; a hint when a local login would have served #463 a Codex row with no session window draws one wide weekly bar Tests #484 #485 #486 the suite asserts behaviour, not the scheduler: wall-clock upper bounds are gone, and subprocess tests spawn the server through `test-helpers/spawn-server.js`, which verifies the server it reached by `server.pid` (new in status) instead of trusting a port Tooling #452 #453 #454 #455 docker workflow actions bumped
Takes KarpelesLab/teamclaude up to ba01b4f while keeping the fork's routing safeguards (403 cooldown without parking, send-failure fail-over, outbound content-length, the single predispatch wait budget, cappedMessage, the overload slot release, the dead-refresh-token guard and refresh retry, the session home wait from d7810f1, and unranked-priority semantics). Adopted from upstream, among others: per-conversation session pins (KarpelesLab#438), 401 fail-over without parking (KarpelesLab#439, KarpelesLab#473), real quota-reset retry-after and candidate counts (KarpelesLab#429, KarpelesLab#408), a headerless 429 retry (KarpelesLab#431), fail-over on a failing 200 stream (KarpelesLab#470), per-account egress proxies (KarpelesLab#441), extra-usage fallback (KarpelesLab#427), maxSpend (KarpelesLab#466), stripOverageHeaders (KarpelesLab#478), keep-alive that outlives the client pool (KarpelesLab#411), a dead terminal not killing the proxy (KarpelesLab#410), console resolution per call (KarpelesLab#432), config reload and sync fixes (KarpelesLab#465, KarpelesLab#415), and the dashboard, TUI and Codex work since 1.1.20. Integration fixes the merge needed beyond conflict hunks: - upstream request-path code that referenced upstream-only locals (sx, route, ctx.tried) rewritten for the fork's forwardRequest - resolveSwitchThreshold was declared twice after a clean auto-merge - the fork's warm-up probe now forwards a routed account through its own proxy instead of sending its credential direct (new regression test) - a routing failure during a token refresh arms the routing hold instead of parking the account; a pinned request may use an account on routing hold - the headerless-429 branch no longer writes after headers were sent - canonical-state allowlists widened for upstream's new quota fields; saves use exportState() - the fork's home wait keys on the conversation pin like selection does Tests adapted where the fork deliberately differs (warm-up probe on by default, account-anchored TUI cursor, coordinator-built Prober and Warmer, unranked priority, per-conversation pin keys, fail-over-only 401, the wait budget instead of inline waits), each with an in-file note. Full suite 2919/2919, lint, typecheck and the strict ratchet (1411 vs 1465) pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- forwardRequest passed a raw account index to five account-manager calls the merge brought in (markCredentialRejected, markRoutingFailed, clearRoutingFailed, clearCredentialRejected, isRoutingDown). An account removed while its request was upstream leaves its successor at that index, so a 401 for the removed account parked the healthy one that moved in. They now pass the account object, which a removal makes stale, not wrong. - The Codex reset-credit read and redemption carried the account's bearer token without its egress proxy (KarpelesLab#441); both now route like every other credentialed call. - A reload that reads a config listing no accounts no longer drops every running account (KarpelesLab#465); it keeps them and logs why. - The MCP set_account_priority description said the default is 0; in this fork an unset priority is unranked and sorts after every number. Each fix has a test that fails without it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This lets one account leave through its own proxy whilst the rest of the fleet carries on as before. The contract is that ALL of that account's traffic goes through the proxy, and ONLY that account's.
{ "name": "waffles@waffle.com", "type": "oauth", "routing": "socks5h://alice:s3cret@proxy.example.com:1080" }upstreamProxymoves every account together and only speaks CONNECT, so it can't do this. Schemes arehttp(CONNECT),socks4,socks4a,socks5andsocks5hwithuser:pass@auth, and thea/hforms resolve at the proxy the way curl's do. There are no new dependencies - the SOCKS handshakes are hand-rolled on the same primitives as the sx CONNECT tunnel, and TLS stays end-to-end in every case.It's a big PR (+3 668 lines, 2 177 of them tests), so here's a map of it, sorry😅
What gets routed
Everything TeamClaude does with the account's credential: forwarding, token refresh, the usage / profile / Codex / backend-quota probes, the OAuth code exchange and profile lookup at login,
teamclaude accountsandteamclaude api. It all goes through oneroutingoption onupstreamFetch/proxyFetch, which outranks both sx andupstreamProxyfor that call. It doesn't chain them - that would be two hops to solve one problem, and the first hop wouldn't be the proxy the operator picked.You can set it from the CLI (
routing <name> [url|none],login/import --routing), the config (accounts[].routing, live on reload throughsync-accounts.js), the TUI (g→ "Account proxy", which sits next to "Upstream proxy") and MCP (set_account_routing). docs/accounts.md has the long version.The URL carries a password, so it's masked everywhere it can be printed: the status payload (attach mode and
status --jsonnever see it), the TUI, a dashboard badge, parse errors, and the MCP write log. That last one needed a smallauditArgshook on the tool definition, because the write log prints tool arguments verbatim.Two things are NOT routed, and the docs say so:
/api/oauth/*,/v1/code/*, its token refresh). That's relayed with the client's credential and belongs to no pooled account, even when it's the same human behind both.login --name <account>fixes it by borrowing the routing the entry already has.When the proxy dies
If you only review one thing closely, make it this.
ECONNREFUSEDsits inSOCKET_TRANSIENT, on the reasoning that every account dials the same host, so failing over is pointless and the right move is to close the connection and let the client retry. From ONE account's proxy that reasoning is backwards. The retry lands on the same account and hits the same dead proxy, forever.So
routingAgentwraps any failure to open a connection through an account's routing (refused, timed out, bad credentials, a target the proxy can't reach, or TLS that never completes over the tunnel) asTEAMCLAUDE_ROUTING_FAILED, and keeps the socket error oncause.isTransientUpstreamErrorreturns false for it, so it takes the ordinary failover path. Nothing of the request has been sent at that point, so failing over can't duplicate anything upstream.The account then sits out for 30 seconds, because otherwise every request queued behind the first one pays the connect timeout before it fails over, and against a black-holed proxy that's up to 30s each. It's modelled on the entitlement cooldown:
routingFailedUntil, unavailable reasonrouting, not probeable while it lasts, and lifted by any response that comes back through the proxy or by changing the account's routing.exhaustedMessagenames it rather than filing it under "quota or rate limit", and theretry-afteris the rest of the hold. A pin still goes to the account it names.Related: a 429 on a routed account no longer arms the sx sticky window or takes the "retrying via sx.org" branch. That 429 was earned by the account's own exit address, and the branch would have re-sent immediately through the proxy that had just been refused.
Testing the proxy before anything depends on it
login --routing,import --routing,routing <name> <url>and the TUI row all open a tunnel to the account's upstream and finish the TLS handshake before they change anything. No request is sent.The reason is OAuth. The code is single-use, so finding a typo'd proxy password at the token exchange costs you the whole browser flow.
--no-checkskips the test for a proxy that isn't up yet, androuting <name> --checkruns it against what's already stored.Also -
--routing=URLworks, a bare--routingis an error rather than being ignored (the alternative is the account being added from your own IP without a word), and--routing nonesigns in direct and clears what's stored.Two TUI fixes that came along for the ride
Both turned up whilst testing the new prompt, and both predate this branch:
_onDatadrops any chunk longer than one character, so a PASTED value vanishes without a sign. Nobody typessocks5h://user:long-password@host:1080by hand, and the sx key and upstream proxy prompts have the same problem.Happy to split those into their own PR if you'd rather take them separately.
Tests
8 new test files, plus additions to the MCP, dashboard, exhausted-message and TUI suites. The full suite (2 260) passes on Node 20, 22, 24 and 25, and lint, typecheck and the strict ratchet (13 fewer than base) are clean. That's all on macOS, so CI will tell us about Linux.
The in-repo tests use SOCKS mocks written next to the client, which means a misreading of the protocol could hide in both. So the branch also went up against real servers - all five schemes against 3proxy, SOCKS5 auth against microsocks, and CONNECT + Basic auth against tinyproxy. They all pass. The handshake bytes for
socks4aandsocks5hwith auth are identical to what curl sends for--socks4aand--socks5-hostname --proxy-user, apart from curl also offering GSSAPI in its greeting.Not done: keep-alive for routed sockets. Every request pays a new tunnel plus a TLS handshake, which is what the
upstreamProxypath does today, and a socket pool felt like too much for a first PR. Per-account VPN egress is next on my list, which is why all of this lives in one module (src/account-routing.js).One thing I'm not precious about is the name.
routingcollides a little with model routes (route,routes, docs/routing.md), which is why the TUI row says "Account proxy" - "Manage routing" was already taken. I keptroutingfor the config key and the CLI because "routing an account's traffic" is how I think about it, butegressorproxywould work just as well. Would you rather I rename it before this goes in?