Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

OPTYM Benchmark

Standardized cost-efficiency benchmark for LLM routing systems.

Compare any LLM proxy, router, or gateway against a fixed set of real evaluation prompts from the most recognized AI benchmarks.

What it measures

Metric Description
Cost savings % saved vs GPT-4o baseline (same prompts, same output)
Latency Avg, p50, p95, p99 response times
Success rate % of requests completed without error
Quality Responses from real benchmark questions — verifiable answers
Throughput Requests per second at given concurrency

Benchmark composition (10,000 unique prompts)

Source Count Category Reference
MMLU 2,000 Knowledge MCQ (57 subjects) Hendrycks et al. 2021
MMLU-Pro 1,500 Hard MCQ (10 choices) Wang et al. 2024
HellaSwag 2,000 Commonsense completion Zellers et al. 2019
GSM8K 1,200 Grade-school math Cobbe et al. 2021
NQ-Open 800 Factual QA Kwiatkowski et al. 2019
WinoGrande 800 Commonsense fill-in-blank Sakaguchi et al. 2020
TriviaQA 700 Trivia questions Joshi et al. 2017
ARC-Challenge 500 Hard science reasoning Clark et al. 2018
ARC-Easy 500 Easy science reasoning Clark et al. 2018

All prompts are unique per run. Cache hit rate should be ~0% to measure real routing efficiency.

Quick start

git clone https://github.com/arturoyo/optym-benchmark.git
cd optym-benchmark

pip install -r requirements.txt

# Run against any OpenAI-compatible endpoint
python benchmark.py \
  --endpoint http://localhost:8081/v1/chat/completions \
  --api-key YOUR_API_KEY \
  --count 100

Scaled runs

# Quick validation
python benchmark.py --count 100

# Medium test
python benchmark.py --count 500

# Full benchmark
python benchmark.py --count 1000

# Maximum (all 10K unique prompts)
python benchmark.py --count 10000 --budget 50.0

Comparing systems

Run the same benchmark against different endpoints:

# OPTY
python benchmark.py --endpoint http://localhost:8081/v1/chat/completions \
  --api-key opty_XXX --count 1000 --label opty

# OpenRouter
python benchmark.py --endpoint https://openrouter.ai/api/v1/chat/completions \
  --api-key sk-or-XXX --count 1000 --model openai/gpt-4o --label openrouter

# Direct GPT-4o (baseline)
python benchmark.py --endpoint https://api.openai.com/v1/chat/completions \
  --api-key sk-XXX --count 1000 --model gpt-4o --label gpt4o-direct

# Compare results
python compare.py results/opty.json results/openrouter.json results/gpt4o-direct.json

Output

Each run produces:

  • results/<label>_results.json — Full metrics
  • results/<label>_quality.json — Sample responses for manual quality review
  • Console report with savings, latency, tier distribution

GPT-4o baseline pricing (March 2026)

Model Input (per 1M tokens) Output (per 1M tokens)
GPT-4o $2.50 $10.00

All savings are calculated relative to what GPT-4o would cost for the same input/output token counts.

What this benchmark is for

This benchmark is designed to compare LLM routers, proxies, and optimization layers — not raw LLM models. The goal is to measure how much a routing/optimization system saves compared to sending everything to GPT-4o.

Compare systems like: OpenRouter, Martian, RouteLLM, Portkey, LiteLLM, OPTY, or any OpenAI-compatible proxy.

Head-to-head: OPTY vs RouteLLM (100 prompts, same benchmark)

Metric OPTY (aggressive) RouteLLM (mf-0.5)
Savings vs GPT-4o 96.7% 94.0%
Total cost $0.0064 $0.0105
Avg latency 8.9s 11.4s
p50 latency 6.8s 9.6s
Success rate 99% 100%
Routing strategy 3-tier (premium 63.6%, mid 35.4%, cheap 1%) Binary (100% gpt-4o-mini)

RouteLLM routes everything to a single weak model. OPTY uses multi-tier, multi-provider intelligent routing — sending hard questions to premium models and easy ones to cheap models.

Sample results: OPTY balanced (1,000 prompts)

Total requests:     1,000
Successful:           995
Savings vs GPT-4o:  41.8%
Cost (OPTY):        $1.01
Cost (GPT-4o):      $1.73
Avg latency:        8.2s
p50 latency:        8.0s
Cache hit rate:     6.2%
Profile:            balanced (default)

Sample results: OPTY aggressive (100 prompts)

Total requests:       100
Successful:            99
Savings vs GPT-4o:  96.7%
Cost (OPTY):        $0.0064
Cost (GPT-4o):      $0.1962
Avg latency:        8.9s
p50 latency:        6.8s
Tier distribution:  premium 63.6% | mid 35.4% | cheap 1.0%
Profile:            aggressive

Run the comparison yourself

# Test your router/proxy
python benchmark.py --endpoint YOUR_ROUTER_URL \
  --api-key YOUR_KEY --model YOUR_MODEL --label my-router --count 100

# Test another router
python benchmark.py --endpoint https://openrouter.ai/api/v1/chat/completions \
  --api-key sk-or-XXX --model openai/gpt-4o --label openrouter --count 100

# Compare results side by side
python compare.py results/my-router_results.json results/openrouter_results.json

Contributing

  1. Fork the repo
  2. Run the benchmark against your system
  3. Submit a PR with your results/<system>_results.json
  4. Include your system name, configuration, and pricing model

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages