English | 简体中文
Stop overpaying for every prompt.
A lightweight local LLM router that automatically balances cost and quality, saving you 90%+ on every session.
| real-world cost reduction | classifier accuracy | request success rate | tests passing |
Built for Codex, Claude Code, Cursor, the OpenAI SDK, and OpenClaw.
"hello" → 🟢 gpt-4.1-nano ($0.0008)
"fix the typo on line 3" → 🟢 deepseek-chat ($0.0012)
"summarize this log in one sentence" → 🟢 minimax-m2.7 ($0.0097)
"refactor this 500-line module" → 🟠 claude-sonnet-4.6 ($0.0337)
"design a distributed scheduler" → 🔴 claude-opus-4.6 ($0.0562)
One endpoint. The router picks the model. You just write code.
Always on premium = burning money on simple tasks. You're on Claude Opus. Every tool call, status check, and log summary in your agent loop bills at Opus prices. One session: $50+. But most of those requests? A $0.001 model handles them fine.
Daily driver on cheap models = stuck when it gets hard. You're on MiniMax or DeepSeek's coding plan. Good enough for everyday work. But when you hit complex architecture or large-scale refactoring, you know Opus should take over — there's just no automatic way to switch.
UncommonRoute fixes exactly this: it automatically picks the right model for every request.
Simple requests go down. Complex requests go up. You don't think about it.
Your client
(Codex / Claude Code / Cursor / OpenAI SDK / OpenClaw)
|
v
UncommonRoute
(runs on your machine)
|
v
Your upstream API
(Commonstack / OpenAI / Ollama / vLLM / Parallax / ...)
| Without UncommonRoute | With UncommonRoute | |
|---|---|---|
| Cost per session | $50+ | ~$3–5 |
| Hard tasks | ✅ Premium model | ✅ Same premium model |
| Simple tasks | ✅ Overkill | ✅ Right-sized |
| Models used | 1 | 15 (auto-selected) |
| Success rate | 28/28 | 28/28 |
Measured in a real Claude Code session.
pip install uncommon-routeuncommon-route route "write a Python function that validates email addresses"
# → shows difficulty tier, model selection, fallback chain# Pick one:
export UNCOMMON_ROUTE_UPSTREAM="https://api.commonstack.ai/v1" # Commonstack
export UNCOMMON_ROUTE_UPSTREAM="https://api.openai.com/v1" # OpenAI
export UNCOMMON_ROUTE_UPSTREAM="http://127.0.0.1:11434/v1" # Ollama / vLLM
export UNCOMMON_ROUTE_API_KEY="your-key-here" # if neededuncommon-route serveThen point your client at the proxy — one line change:
| Client | Change this |
|---|---|
| Codex / Cursor / OpenAI SDK | export OPENAI_BASE_URL="http://localhost:8403/v1" |
| Claude Code | export ANTHROPIC_BASE_URL="http://localhost:8403" |
| OpenClaw | Plugin — see OpenClaw integration |
Done. Your existing workflow is already saving money.
uncommon-route doctor # one command checks everythingThe classifier estimates a difficulty score (0.0–1.0) from structural features and character n-grams. No keyword lists, no hardcoded rules. SIMPLE / MEDIUM / COMPLEX appear in logs and the dashboard, but they're display labels, not decision boundaries.
| Mode | Virtual model ID | Behavior |
|---|---|---|
| auto | uncommon-route/auto |
Balanced — best quality-per-dollar, adapts with difficulty |
| fast | uncommon-route/fast |
Cost-first — cheapest acceptable model |
| best | uncommon-route/best |
Quality-first — highest quality, cost nearly ignored |
Only virtual model IDs trigger routing. Explicit real model IDs pass through unchanged.
Model quality comes from PinchBench agent task scores, not price assumptions. The selector uses Thompson Sampling (one Beta distribution per model) — models with fewer observations get wider distributions and chances to prove themselves. Quality scores improve over time through Bayesian updating.
| Layer | Source | What it learns |
|---|---|---|
| Benchmark prior | PinchBench API + seed data | Model quality baselines |
| Implicit feedback | HTTP failures, retrial detection, logprob confidence | Automatic quality signals per request |
| Explicit feedback | User ok / weak / strong signals | Direct corrections — 3 clicks to change routing |
Tools in the request body don't inflate difficulty. A "hello" through Claude Code still routes as SIMPLE. The classifier evaluates the user's prompt on its own structural merits.
uncommon-route serve
# open http://127.0.0.1:8403/dashboard/See request counts, latency, cost savings, model distribution, selector state, spend limits, and recent feedback in real time.
uncommon-route serve --daemon # background mode
uncommon-route stop
uncommon-route logs --follow
uncommon-route stats| Variable | Meaning |
|---|---|
UNCOMMON_ROUTE_UPSTREAM |
Upstream OpenAI-compatible API URL |
UNCOMMON_ROUTE_API_KEY |
API key for the upstream provider |
UNCOMMON_ROUTE_PORT |
Local proxy port (default 8403) |
uncommon-route provider add openai sk-your-openai-key
uncommon-route provider add anthropic sk-ant-your-key
uncommon-route provider list
uncommon-route provider modelsuncommon-route config set-default-mode fast
uncommon-route config set-tier auto SIMPLE moonshot/kimi-k2.5 \
--fallback google/gemini-2.5-flash-lite,deepseek/deepseek-chat
uncommon-route config set-tier best COMPLEX anthropic/claude-opus-4.6 \
--fallback anthropic/claude-sonnet-4.6 --strategy hard-pinNote: The live pool scorer runs at request time and does not yet enforce
--strategy hard-pinat the request level. To force a specific model immediately, send that non-virtual model ID directly.
uncommon-route spend set per_request 0.10
uncommon-route spend set hourly 5.00
uncommon-route spend set daily 20.00
uncommon-route spend statusReturns HTTP 429 with reset_in_seconds when a limit is hit.
| Client type | Base URL |
|---|---|
| OpenAI-compatible | http://127.0.0.1:8403/v1 |
| Anthropic-style | http://127.0.0.1:8403 |
| Endpoint | Purpose |
|---|---|
GET /health |
Liveness + config status |
GET /v1/models |
Virtual models exposed by the router |
GET /v1/models/mapping |
Internal-to-upstream model mapping |
GET /v1/selector |
Inspect or preview routing decisions |
POST /v1/feedback |
Submit quality feedback |
GET /dashboard/ |
Monitoring UI |
x-uncommon-route-model · x-uncommon-route-tier · x-uncommon-route-mode · x-uncommon-route-reasoning
from uncommon_route import classify, route
decision = route("explain the Byzantine Generals Problem")
print(decision.model) # "anthropic/claude-sonnet-4.6"
print(decision.tier) # "COMPLEX"
print(decision.confidence) # 0.87Full API reference: docs/api.md
Composition pipeline — handling oversized tool outputs
The proxy can compact oversized text/JSON, offload large tool results to local artifacts, create semantic side-channel summaries, and checkpoint long histories. Artifacts are stored in ~/.uncommon-route/artifacts/.
Anthropic-native transport
When routing lands on an Anthropic-family model, UncommonRoute can preserve Anthropic-native transport and caching semantics while serving OpenAI-style clients normally.
Local classifier retraining
The classifier uses structural features and character n-grams only — no keyword lists. Retrain on your own data:
python -c "from uncommon_route.router.classifier import train_and_save_model; train_and_save_model('bench/data/train.jsonl')"Model discovery and mapping
UncommonRoute fetches /v1/models from your upstream, builds a live model pool, maps internal IDs to what the upstream actually serves, and records learned aliases when fallbacks find a better match.
uncommon-route doctor
curl http://127.0.0.1:8403/v1/models/mapping1,904 training samples, 1,077 held-out test samples:
| Metric | Value |
|---|---|
| Training accuracy | 99.2% |
| Held-out accuracy | 88.5% |
The classifier provides a continuous difficulty signal. Benchmark quality data and Thompson Sampling compensate for classification noise.
End-to-end testing through Claude Code with Commonstack upstream:
| Metric | Value |
|---|---|
| Cost reduction | ~90–95% (vs always-premium) |
| Request success rate | 28/28 |
| Models auto-selected | 15 |
| Expensive model waste on simple tasks | 0 |
| Clicks to change routing | 3 |
python -m bench.run # reproduce it yourself| Symptom | Fix |
|---|---|
route works but real requests fail |
Check UNCOMMON_ROUTE_UPSTREAM and UNCOMMON_ROUTE_API_KEY, run uncommon-route doctor |
| Codex / Cursor can't connect | OPENAI_BASE_URL must end with /v1 |
| Claude Code can't connect | ANTHROPIC_BASE_URL should point to the router root, not /v1 |
| Local upstream discovery fails | Some servers have /chat/completions but no /models; passthrough may work; doctor will tell you |
| Don't know what to run first | uncommon-route doctor |
# Stop the proxy
uncommon-route stop
# Remove local state (stats, feedback, learning weights)
rm -rf "${UNCOMMON_ROUTE_DATA_DIR:-$HOME/.uncommon-route}"
# Restore client config
unset OPENAI_BASE_URL ANTHROPIC_BASE_URL UNCOMMON_ROUTE_UPSTREAM UNCOMMON_ROUTE_API_KEY
# Uninstall
pip uninstall uncommon-routeIf you installed the OpenClaw plugin: openclaw plugins uninstall @anjieyang/uncommon-route
| Directory | Contents |
|---|---|
uncommon_route/ |
Shipped runtime: proxy, router, CLI, calibration |
bench/ |
Offline evaluation datasets and benchmark scripts |
demo/ |
Local comparison / demo apps |
frontend/ |
Dashboard and demo frontends |
git clone https://github.com/CommonstackAI/UncommonRoute.git
cd UncommonRoute
pip install -e ".[dev]"
python -m pytest tests -v # 341 tests passingMIT — see LICENSE.