Skip to content

feat(quota): read the Z.ai GLM Coding Plan windows for a z.ai backend account - #494

Merged
MagicalTux merged 4 commits into
KarpelesLab:masterfrom
thomasahle:pr/zai-quota
Oct 3, 2026
Merged

MagicalTux merged 4 commits into
KarpelesLab:masterfrom
thomasahle:pr/zai-quota

Conversation

@thomasahle

Copy link
Copy Markdown
Contributor

Summary

A z.ai account already works as a third-party backend (upstream + modelMap), but its row reads unknown: the GLM Coding Plan meters a five-hour and a weekly token window, and nothing read them. This adds the provider entry, so teamclaude status shows:

  z.ai (oauth, prio 100) active
  Plan     [██░░░░░░░░░░░░░░░░] 12% · 5h 12% (resets 2h10m) · week 37% (resets 3d4h)

What it does

  • Provider entries for api.z.ai and open.bigmodel.cn (the mainland plan; same monitor path, same reply) in src/backend-quota.js. Path /api/monitor/usage/quota/limit, resolved against the account's upstream like DeepSeek's.
  • Parses the TOKENS_LIMIT rows — unit: 3, number: 5 is the five-hour window, unit: 6, number: 1 the weekly — each with percentage (0-100, used) and nextResetTime (ms). Both windows go into one normalized reading: utilization is the fuller of the two (the one that decides whether the account can serve), the text names each with its reset. The TIME_LIMIT row is the monthly MCP-tool allowance, not a token quota, and is left out. Percentages are clamped to 0-100; a missing reset drops only the countdown.
  • Small extension to the module: z.ai's monitor wants the raw key in Authorization, not Bearer <key> (the chat endpoint on the same host accepts either). A provider may now supply headers(credential); every other provider keeps the bearer default, and the prober, the quota field and the renderer still never name a provider — the design the module's header comment sets out.

Test plan

  • npm test — 2589 pass. New test/backend-quota-zai.test.js (7): host matching incl. a look-alike host; URL and raw-key header; the combined reading; MCP-only or unrecognized replies → unrecognized response; clamping and missing reset; HTTP failure; DeepSeek still sends a bearer.
  • npm run lint, npm run typecheck
  • Against a live z.ai plan. The reply shape and the raw-key requirement come from the endpoint's use in several open-source quota readers (glm-quota, meterbar, ThinkWatch), not from z.ai's documentation, which does not describe the monitor endpoint. I have a plan and will report the live reading here once its key is in my fleet.

Docs: the provider list in docs/quota.md; a sample z.ai account with a modelMap in docs/accounts.md.

🤖 Generated with Claude Code

… account

A z.ai account already works as a third-party Anthropic-compatible
backend (upstream + modelMap). What it lacked was any reading of its
own quota: the plan meters a five-hour and a weekly token window, and
the row read `unknown`.

Z.ai publishes both at /api/monitor/usage/quota/limit on the same host
as the chat endpoint, as used-percentages with a reset timestamp each.
This adds a provider entry for `api.z.ai` and for `open.bigmodel.cn`
(the mainland plan, same reply): both windows go into one normalized
reading — the bar is the fuller of the two, the one that decides
whether the account can serve, and the text names each with its reset,
`5h 12% (resets 2h10m) · week 37% (resets 3d4h)`. The monthly MCP-tool
row is not a token quota and is left out.

One quirk needed a small extension: the monitor wants the raw key in
`Authorization`, not `Bearer <key>` (the chat endpoint accepts either).
A provider may now name its own header shape; every other backend keeps
the bearer default, and the prober, the quota field and the renderer
still never name a provider.

Docs: the provider list in quota.md, and a sample z.ai account with a
modelMap in accounts.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ity does

The z.ai sample suggested leaving priority off and adding a glm-* route to
reserve the account for explicit GLM sessions. A route only restricts who
serves the models it matches; every other model can still be served by
the account, so without a high priority it would take ordinary traffic,
with the modelMap rewriting Claude names onto GLM. Say so.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@thomasahle
thomasahle marked this pull request as draft September 30, 2026 08:22
MagicalTux and others added 2 commits October 3, 2026 12:06
… passes the strict ratchet

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@MagicalTux
MagicalTux marked this pull request as ready for review October 3, 2026 06:59
@MagicalTux

Copy link
Copy Markdown
Member

Reviewed and merging. Rebased onto #495 so the two readers share one formatUntil, and gave the provider table a BackendQuotaProvider typedef so the optional headers() passes the strict typecheck ratchet (the node 24 failure on the original run). Marked ready on your behalf for the same reason as on #495.

@MagicalTux
MagicalTux merged commit 7d76f92 into KarpelesLab:master Oct 3, 2026
5 checks passed
@MagicalTux MagicalTux mentioned this pull request Oct 4, 2026
MagicalTux added a commit that referenced this pull request Oct 4, 2026
Fifteen commits since 1.1.22. Two change what an existing install does
without an opt-in (#514, #503); the rest is additive or display.

Behaviour changes
  #514 the TUI quota-bar percentage (`quotaBarPercent`) is off unless set:
       a bar carries its countdown, and its fill is the percentage. The
       switch (g → Bar percentage) is unchanged; a config without the key
       now reads as off
  #503 with Codex accounts in the pool, a request on the intercepted
       chatgpt.com that is not Codex inference (codex-cli 0.156's workspace
       discovery, plugin, MCP and settings calls) is passed through to
       chatgpt.com with the client's own login instead of reaching the
       Anthropic pool and a 404; the Codex Responses WebSocket is refused
       (501) so the CLI falls back to HTTPS, where the pool serves it (#492)
  #502 `login --api` adds a key at priority 100, a last resort behind the
       subscriptions, with `--priority <n>` to place it; existing entries
       are untouched (#497)

Fixes
  #498 artifacts (`/api/frame/*`) are relayed with the client's own
       credential, so publishing and reading them works behind a rotated pool
  #499 the Remote Control bridge (`/v1/environments/*`, `/v1/sessions/*`,
       `/v2/session_ingress/*`, `/v2/ccr-sessions/*`), which Claude Code
       2.1.287 sends through HTTPS_PROXY, is relayed with the client's own
       credential; registration no longer lands in a rotated account's org
  #488 keep-warm works on Windows: the warm-up client is spawned through a
       shell, so npm's `claude.cmd` shim is found
  #489 Node 24.6's undefined HTTP/2 keep-alive buffer no longer crashes the
       MITM tunnel with a NaN timeout
  #490 on a very wide terminal the two provider panes sit together and their
       bars grow to the list's cap instead of padding half the screen

Features
  #496 `accountSort` orders the TUI account list by the soonest reset of a
       window (session, weekly, S7, F7) inside each provider group; cycled
       from the settings screen, display only
  #494 a z.ai GLM Coding Plan backend account shows its 5-hour and weekly
       windows in `teamclaude status`
  #495 a NanoGPT backend account shows its daily and weekly windows and
       NanoGPT's own billing advice (`billing balance`, `balance not allowed`)
  #500 each dashboard account card lists the models the account served in
       the last 15 minutes

Docs
  #505 what a cross-organization switch costs a Sonnet 5.5 conversation:
       the API drops the earlier thinking blocks silently; same-org pools are
       unaffected (#491)

Tests
  #501 #504 two tests that raced the wall clock or a shared port now assert
       the mechanism, and the last private server-spawn harness is gone

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants