Models change every month. Your harness shouldn't.
install · quick start · models · docs · remote daemon · islands
One agent harness, every model. llmux is a local Anthropic-compatible proxy for Claude Code: claude talks to http://localhost:3456, llmux decides which account/backend serves the request. Your subagents, slash commands, MCP servers, hooks, and CLAUDE.md conventions stay put while frontier models and subscription limits keep moving — /model fable, /model gpt-5.6-sol, /model grok-4.7 are routing signals, not migrations.
- one Rust binary — daemon, live TUI dashboard, login/import, updater, and a Claude Code launcher (
llmux run) - four backend groups in one pool — Claude (subscription + API key), Codex (
gpt-*/ ChatGPT), Grok (grok-*/ xAI), OpenRouter (or-*/ free models on an OpenRouter key), routed by model name (models →) - multi-account scheduling — quota-aware perishability scoring or sticky round-robin, 429 cooldown parking, Fable weekly ceilings (schedulers →)
- DevTools for your agent's model traffic — live per-request receipts, a raw request/response viewer over all four wire legs, copy-as-curl (the accidental AI debugger →)
- remote-first — one central daemon, every other machine a pure client, with per-machine multi-tenant keys (remote daemon →)
- llmux Islands — native macOS menu-bar/notch companion, plus a KDE/Qt port (islands →)
The bet behind it — the model is a consumable, the harness is capital — is in why llmux exists. The complete feature list lives in what ships today.
brew install 2lab-ai/tap/llmuxRolling preview channel:
brew install 2lab-ai/tap/llmux-previewOptional native macOS companion (the KDE port is a source build):
brew install 2lab-ai/tap/llmux-islandsBuild from source:
git clone https://github.com/2lab-ai/llmux && cd llmux
just build # cargo build --release --lockedAdd accounts:
llmux login # Claude subscription OAuth; repeat once per account
llmux login --api # optional: Anthropic API key
llmux login --codex # optional: Codex / ChatGPT subscription
llmux login --grok # optional: Grok / xAI (device-code flow)
llmux login --openrouter # optional: OpenRouter (browser PKCE; --paste for an existing key)
llmux import # or import supported local credential storesAlready looking at the dashboard? n opens a provider picker for the same four browser logins (Claude / Codex / Grok / OpenRouter) — the flow runs in the client and the credential is injected into the daemon, so it works attached to a remote one too.
Run Claude Code through llmux:
llmux run # starts/reuses the daemon, then launches claude
alias lx='llmux run' # a convenient alias; args after -- pass through to claudeInside that session /model lists the llmux catalog — every codex/grok/openrouter id too, not just the built-in Claude rows. The same launch exports ANTHROPIC_DEFAULT_{OPUS,FABLE,SONNET,HAIKU}_MODEL from the catalog's alias owners, so /model opus — which Claude Code resolves natively, before llmux ever sees it — lands on claude-opus-5-5[1m] and its 1M window instead of the client's 200k default (alias exports; a var you already export is left alone). --no-model-picker opts out of both.
Want the foreground TUI dashboard instead:
llmux serverManual shell wiring also works: eval "$(llmux env)", then claude.
Claude Code's model name becomes the routing signal:
/model fable
/model opus[1m]
/model gpt-5.6-sol[1m]
/model grok-4.7
/model or-ox-alpha
| Name pattern | Backend group |
|---|---|
Claude-like (fable, opus, sonnet, haiku, claude-*) |
Claude accounts |
gpt-* / codex / aliases (sol, terra, luna) |
Codex accounts |
grok / grok-* |
Grok accounts |
or / or-* / openrouter/* |
OpenRouter accounts |
Curated catalog (ids, aliases, efforts, context windows): GET /models and docs/models.md. Routing config: docs/configuration.md.
Same request, different backend — read provider compatibility before you trust a field. Claude and OpenRouter are passthrough; Codex and Grok are subscription gateways llmux translates onto, and they do not honor everything Claude Code sends.
gpt-*(Codex): no output-limit guarantee — yourmax_tokensis not sent upstream at all. The gateway answered400 Unsupported parameter: max_output_tokens(live probe 2026-09-14), and no supported alternative cap field was found in the current official Codex client or its docs (read 2026-09-14), so llmux omits the cap rather than faking one. That is a search result, not an allowlist: other field names are untested, not proven absent.grok-*: the cap is forwarded, but it is not the budget you asked for. Amax_output_tokens: 1probe (2026-09-14) came backincompletewith one visible token and 168 reported output tokens, 167 of them reasoning. What it bounds in general — and what it costs — is unmeasured.- Both: non-null
temperature/top_p/top_kand non-emptystop_sequencesare refused with a local 400, priorthinkingblocks are dropped, and there is no reasoning continuity across turns.A translated response that lost something names it in
X-Llmux-Omitted-Fields/X-Llmux-Compatibility-Warnings(a faithful one carries neither header); sendX-Llmux-Compatibility: strictto turn any such loss into a 400 instead. Full matrix, receipts and known unknowns: docs/provider-compatibility.md.
llmux channel # print the current channel (stable | preview)
llmux update # upgrade in place; restarts the daemon only if the binary changed
llmux channel preview # switch channels (mirrored onto the llmux-islands cask)Details: channels and updating.
- docs index — map of all guides
- why llmux exists — the harness-is-capital bet
- what ships today — the complete feature list
- the accidental AI debugger — per-request receipts, raw request/response viewer, copy-as-curl, email masking
- remote daemon — one central daemon, remote-mode command matrix, transport security
- schedulers — eligibility gates,
defaultvsround-robin, adding a mode - operational reference — commands, TUI keys, daemon/dashboard, multi-tenant keys
- configuration — config keys, proxy/scheduler/routing, account types
- models — catalog, aliases, context windows, group routing
- provider compatibility — per-backend difference matrix: dropped/refused request fields,
max_tokenson Codex/Grok, diagnostic headers - FAQ — context-window workarounds (
gpt-*→ Claude 1M/compact→ back) - llmux Islands — macOS menu-bar/notch companion
- system prompts (multi-model) — real captured wire system prompts
llmux is for one human using their own accounts — no credential pooling, no resale.
- Durable path: Claude Code as the harness; Claude through Claude Code/subscription or Anthropic API keys; other models through supported API keys.
- Convenience path: routing third-party flat-rate subscription tokens through Claude Code depends on that vendor's current policy and can change without notice. Use it opt-in, with your own accounts only, and keep an API-key fallback configured.
- Anthropic quota headers and vendor subscription-token behavior may change.
- llmux is not affiliated with Anthropic, OpenAI, xAI, or OpenRouter.
Product intent — what llmux is, what it bets on, and what it refuses — is fixed in .prd/.
If you are an AI agent working on this repository, read AGENTS.md before making changes.
MIT.
