Repository navigation
feat(quota): read the Z.ai GLM Coding Plan windows for a z.ai backend account - #494
Merged
Merged
Conversation
… account A z.ai account already works as a third-party Anthropic-compatible backend (upstream + modelMap). What it lacked was any reading of its own quota: the plan meters a five-hour and a weekly token window, and the row read `unknown`. Z.ai publishes both at /api/monitor/usage/quota/limit on the same host as the chat endpoint, as used-percentages with a reset timestamp each. This adds a provider entry for `api.z.ai` and for `open.bigmodel.cn` (the mainland plan, same reply): both windows go into one normalized reading — the bar is the fuller of the two, the one that decides whether the account can serve, and the text names each with its reset, `5h 12% (resets 2h10m) · week 37% (resets 3d4h)`. The monthly MCP-tool row is not a token quota and is left out. One quirk needed a small extension: the monitor wants the raw key in `Authorization`, not `Bearer <key>` (the chat endpoint accepts either). A provider may now name its own header shape; every other backend keeps the bearer default, and the prober, the quota field and the renderer still never name a provider. Docs: the provider list in quota.md, and a sample z.ai account with a modelMap in accounts.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2 of 3 tasks
…ity does The z.ai sample suggested leaving priority off and adding a glm-* route to reserve the account for explicit GLM sessions. A route only restricts who serves the models it matches; every other model can still be served by the account, so without a high priority it would take ordinary traffic, with the modelMap rewriting Claude names onto GLM. Say so. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
thomasahle
marked this pull request as draft
September 30, 2026 08:22
… passes the strict ratchet Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e NanoGPT reader from KarpelesLab#495)
MagicalTux
marked this pull request as ready for review
October 3, 2026 06:59
Member
|
Reviewed and merging. Rebased onto #495 so the two readers share one |
Merged
MagicalTux
added a commit
that referenced
this pull request
Oct 4, 2026
Fifteen commits since 1.1.22. Two change what an existing install does without an opt-in (#514, #503); the rest is additive or display. Behaviour changes #514 the TUI quota-bar percentage (`quotaBarPercent`) is off unless set: a bar carries its countdown, and its fill is the percentage. The switch (g → Bar percentage) is unchanged; a config without the key now reads as off #503 with Codex accounts in the pool, a request on the intercepted chatgpt.com that is not Codex inference (codex-cli 0.156's workspace discovery, plugin, MCP and settings calls) is passed through to chatgpt.com with the client's own login instead of reaching the Anthropic pool and a 404; the Codex Responses WebSocket is refused (501) so the CLI falls back to HTTPS, where the pool serves it (#492) #502 `login --api` adds a key at priority 100, a last resort behind the subscriptions, with `--priority <n>` to place it; existing entries are untouched (#497) Fixes #498 artifacts (`/api/frame/*`) are relayed with the client's own credential, so publishing and reading them works behind a rotated pool #499 the Remote Control bridge (`/v1/environments/*`, `/v1/sessions/*`, `/v2/session_ingress/*`, `/v2/ccr-sessions/*`), which Claude Code 2.1.287 sends through HTTPS_PROXY, is relayed with the client's own credential; registration no longer lands in a rotated account's org #488 keep-warm works on Windows: the warm-up client is spawned through a shell, so npm's `claude.cmd` shim is found #489 Node 24.6's undefined HTTP/2 keep-alive buffer no longer crashes the MITM tunnel with a NaN timeout #490 on a very wide terminal the two provider panes sit together and their bars grow to the list's cap instead of padding half the screen Features #496 `accountSort` orders the TUI account list by the soonest reset of a window (session, weekly, S7, F7) inside each provider group; cycled from the settings screen, display only #494 a z.ai GLM Coding Plan backend account shows its 5-hour and weekly windows in `teamclaude status` #495 a NanoGPT backend account shows its daily and weekly windows and NanoGPT's own billing advice (`billing balance`, `balance not allowed`) #500 each dashboard account card lists the models the account served in the last 15 minutes Docs #505 what a cross-organization switch costs a Sonnet 5.5 conversation: the API drops the earlier thinking blocks silently; same-org pools are unaffected (#491) Tests #501 #504 two tests that raced the wall clock or a shared port now assert the mechanism, and the last private server-spawn harness is gone Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A z.ai account already works as a third-party backend (
upstream+modelMap), but its row readsunknown: the GLM Coding Plan meters a five-hour and a weekly token window, and nothing read them. This adds the provider entry, soteamclaude statusshows:What it does
api.z.aiandopen.bigmodel.cn(the mainland plan; same monitor path, same reply) insrc/backend-quota.js. Path/api/monitor/usage/quota/limit, resolved against the account'supstreamlike DeepSeek's.TOKENS_LIMITrows —unit: 3, number: 5is the five-hour window,unit: 6, number: 1the weekly — each withpercentage(0-100, used) andnextResetTime(ms). Both windows go into one normalized reading:utilizationis the fuller of the two (the one that decides whether the account can serve), the text names each with its reset. TheTIME_LIMITrow is the monthly MCP-tool allowance, not a token quota, and is left out. Percentages are clamped to 0-100; a missing reset drops only the countdown.Authorization, notBearer <key>(the chat endpoint on the same host accepts either). A provider may now supplyheaders(credential); every other provider keeps the bearer default, and the prober, the quota field and the renderer still never name a provider — the design the module's header comment sets out.Test plan
npm test— 2589 pass. Newtest/backend-quota-zai.test.js(7): host matching incl. a look-alike host; URL and raw-key header; the combined reading; MCP-only or unrecognized replies →unrecognized response; clamping and missing reset; HTTP failure; DeepSeek still sends a bearer.npm run lint,npm run typecheckDocs: the provider list in
docs/quota.md; a sample z.ai account with amodelMapindocs/accounts.md.🤖 Generated with Claude Code