Standardized cost-efficiency benchmark for LLM routing systems.
Compare any LLM proxy, router, or gateway against a fixed set of real evaluation prompts from the most recognized AI benchmarks.
| Metric | Description |
|---|---|
| Cost savings | % saved vs GPT-4o baseline (same prompts, same output) |
| Latency | Avg, p50, p95, p99 response times |
| Success rate | % of requests completed without error |
| Quality | Responses from real benchmark questions — verifiable answers |
| Throughput | Requests per second at given concurrency |
| Source | Count | Category | Reference |
|---|---|---|---|
| MMLU | 2,000 | Knowledge MCQ (57 subjects) | Hendrycks et al. 2021 |
| MMLU-Pro | 1,500 | Hard MCQ (10 choices) | Wang et al. 2024 |
| HellaSwag | 2,000 | Commonsense completion | Zellers et al. 2019 |
| GSM8K | 1,200 | Grade-school math | Cobbe et al. 2021 |
| NQ-Open | 800 | Factual QA | Kwiatkowski et al. 2019 |
| WinoGrande | 800 | Commonsense fill-in-blank | Sakaguchi et al. 2020 |
| TriviaQA | 700 | Trivia questions | Joshi et al. 2017 |
| ARC-Challenge | 500 | Hard science reasoning | Clark et al. 2018 |
| ARC-Easy | 500 | Easy science reasoning | Clark et al. 2018 |
All prompts are unique per run. Cache hit rate should be ~0% to measure real routing efficiency.
git clone https://github.com/arturoyo/optym-benchmark.git
cd optym-benchmark
pip install -r requirements.txt
# Run against any OpenAI-compatible endpoint
python benchmark.py \
--endpoint http://localhost:8081/v1/chat/completions \
--api-key YOUR_API_KEY \
--count 100# Quick validation
python benchmark.py --count 100
# Medium test
python benchmark.py --count 500
# Full benchmark
python benchmark.py --count 1000
# Maximum (all 10K unique prompts)
python benchmark.py --count 10000 --budget 50.0Run the same benchmark against different endpoints:
# OPTY
python benchmark.py --endpoint http://localhost:8081/v1/chat/completions \
--api-key opty_XXX --count 1000 --label opty
# OpenRouter
python benchmark.py --endpoint https://openrouter.ai/api/v1/chat/completions \
--api-key sk-or-XXX --count 1000 --model openai/gpt-4o --label openrouter
# Direct GPT-4o (baseline)
python benchmark.py --endpoint https://api.openai.com/v1/chat/completions \
--api-key sk-XXX --count 1000 --model gpt-4o --label gpt4o-direct
# Compare results
python compare.py results/opty.json results/openrouter.json results/gpt4o-direct.jsonEach run produces:
results/<label>_results.json— Full metricsresults/<label>_quality.json— Sample responses for manual quality review- Console report with savings, latency, tier distribution
| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| GPT-4o | $2.50 | $10.00 |
All savings are calculated relative to what GPT-4o would cost for the same input/output token counts.
This benchmark is designed to compare LLM routers, proxies, and optimization layers — not raw LLM models. The goal is to measure how much a routing/optimization system saves compared to sending everything to GPT-4o.
Compare systems like: OpenRouter, Martian, RouteLLM, Portkey, LiteLLM, OPTY, or any OpenAI-compatible proxy.
| Metric | OPTY (aggressive) | RouteLLM (mf-0.5) |
|---|---|---|
| Savings vs GPT-4o | 96.7% | 94.0% |
| Total cost | $0.0064 | $0.0105 |
| Avg latency | 8.9s | 11.4s |
| p50 latency | 6.8s | 9.6s |
| Success rate | 99% | 100% |
| Routing strategy | 3-tier (premium 63.6%, mid 35.4%, cheap 1%) | Binary (100% gpt-4o-mini) |
RouteLLM routes everything to a single weak model. OPTY uses multi-tier, multi-provider intelligent routing — sending hard questions to premium models and easy ones to cheap models.
Total requests: 1,000
Successful: 995
Savings vs GPT-4o: 41.8%
Cost (OPTY): $1.01
Cost (GPT-4o): $1.73
Avg latency: 8.2s
p50 latency: 8.0s
Cache hit rate: 6.2%
Profile: balanced (default)
Total requests: 100
Successful: 99
Savings vs GPT-4o: 96.7%
Cost (OPTY): $0.0064
Cost (GPT-4o): $0.1962
Avg latency: 8.9s
p50 latency: 6.8s
Tier distribution: premium 63.6% | mid 35.4% | cheap 1.0%
Profile: aggressive
# Test your router/proxy
python benchmark.py --endpoint YOUR_ROUTER_URL \
--api-key YOUR_KEY --model YOUR_MODEL --label my-router --count 100
# Test another router
python benchmark.py --endpoint https://openrouter.ai/api/v1/chat/completions \
--api-key sk-or-XXX --model openai/gpt-4o --label openrouter --count 100
# Compare results side by side
python compare.py results/my-router_results.json results/openrouter_results.json- Fork the repo
- Run the benchmark against your system
- Submit a PR with your
results/<system>_results.json - Include your system name, configuration, and pricing model
MIT