Open-source, OpenAI-compatible LLM gateway you run yourself. One endpoint for 40+ providers, with virtual keys, budgets, and usage tracking.
-
Updated
Sep 4, 2026 - Python
Open-source, OpenAI-compatible LLM gateway you run yourself. One endpoint for 40+ providers, with virtual keys, budgets, and usage tracking.
RouterArena: An open framework for evaluating LLM routers with standardized datasets, metrics, an automated framework, and a live leaderboard.
A prompt-aware LLM router that predicts which models can complete each request, then selects the cheapest capable one: 53.2% lower cost and +1.9 pts completion on our tested dataset.
Combine the results from a panel of models into an enhanced response
Stop overpaying to run your agents. Kalibr routes every request to lower-cost model and tool paths without degrading performance.
Automatic cost-aware model routing plugin for Hermes Agent
Per-step LLM routing benchmark with 970 static labels, live SWE-bench evaluation, an open data pipeline, and a public leaderboard.
Skill + Agent + Model + Thinking depth — auto-routed before any tool fires. One SKILL.md for Claude Code. 90% routing accuracy, per-step model enforcement, 30%+ savings on multi-step chains.
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface
Budget-Aware Agentic Routing (BAAR) — Intelligent LLM model selection with a zero-call financial kill-switch. Save 90% on costs without losing accuracy.
[ACL 2026] Official implementation of MTRouter, a cost-aware multi-turn LLM routing framework accepted to ACL 2026 Main Conference.
[ICML2026] The official code of "Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale"
Free LLM router - latency-based routing across 31 NVIDIA NIM models with automatic failover.
Use Codex from Claude Code, or Claude Code from Codex, through the subscriptions you already have. Multi-turn consultations between agent CLIs with no provider API key, a validated response envelope, and a local dashboard of every consultation.
Automatic model+effort routing for Claude Code: tiered subagents, escalation ladder, independent verifier
Open-source LLM cost observability and smart routing that runs in-process. No proxy, no telemetry, zero added latency.
Interactive 3D knowledge graph of 793 open-source AI harnesses — governance, agent, evaluation, red-team, routing — plus free education. 97% curator-reviewed (incl. a three-model ensemble pass). Built after the 2026-06-13 Fable/Mythos worldwide recall proved closed-weights models can vanish overnight.
3-Tier hybrid AI router that orchestrates FunctionGemma-270M on-device and Gemini 2.5 Flash Lite in the cloud for 99% function-calling accuracy at 548ms avg latency. Built at the Cactus × Google DeepMind Hackathon.
Privacy-aware, local-first router for your CLI coding agents (Codex, Claude Code) and local LLMs (Ollama) — keeps sensitive prompts on-device and cuts premium-model usage.
Risk-adaptive Codex workflow with policy-level caps, native quality gates, and optional Superpowers handoff.
To associate your repository with the llm-routing topic, visit your repo's landing page and select "manage topics."