Skip to content

About

Measured token-efficiency statistics for the Claude API: 82 real calls, frozen payloads, raw data, reproducible analysis.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

🔬 claude-token-efficiency

analyze python license data

Real, measured token-efficiency statistics for the Claude API — 82 live API calls, frozen payloads, raw data, reproducible analysis. Every number in this repository comes from the usage field of an actual Anthropic Messages API response; nothing is estimated or simulated.

English first · Bahasa Indonesia di bawah 🇮🇩


⚠️ Which model was actually measured?

The frozen plan requested claude-fable-5. The claude.ai artifact channel that executed this run silently served claude-sonnet-4-6 instead — this is recorded in the model field of every one of the 82 API responses, which is the source of truth (see results/results.json). The harness's own meta.model_used reflects only what was requested, so the per-record field wins.

All statistics below therefore describe claude-sonnet-4-6 as measured on 2026-08-01 through that channel. The repo ships two reproduction paths (browser harness and CLI) that anyone with an API key can point at claude-fable-5 or any other Claude model — the plan, analysis, and charts regenerate automatically.

One more channel artifact worth knowing: a message containing a single character measured 4,173 input tokens, i.e. the claude.ai channel injects a ≈ 4,172-token envelope (its own system prompt) into every call. All input-side tables below subtract this measured envelope; the subtraction is validated by duplicate payloads counting identically (see Methodology).

Key findings (measured, not estimated)

# Finding Measured numbers
1 CSV/TSV is the cheapest way to hand Claude tabular data; XML is the most expensive. Same 25 records: csv/tsv 618 tokens, md-table 853, minified JSON 958, YAML 1,155, pretty JSON 1,659, XML 1,963 → XML costs 3.18× CSV; switching XML→CSV saves 68.5%.
2 Pretty-printing JSON costs +73%; indenting deeper costs nothing. Minified 958 → indent-2 1,659 (+73.2%); indent-4 and indent-8 also exactly 1,659 despite +2,100 chars of extra spaces — whitespace runs collapse into the same token count.
3 Per-record marginal cost: csv 24.3, markdown table 33.3, pretty JSON 66.3 tokens/record (linear, r² = 1.0000). JSON's key names are re-paid on every record; CSV pays them once in the header.
4 A one-line brevity instruction cut output by ~73%. Pooled over 3 questions × 3 reps: no instruction 349.4 ± 89.3 output tokens → "answer in at most 2 sentences" 94.6 ± 27.2 (−72.9%); concise system prompt 127.8 ± 84.0 (−63.4%). User-turn instruction beat the system prompt here.
5 Indonesian costs more tokens than English on both sides. Input: 2.90 ± 0.11 chars/token (ID) vs 4.76 ± 0.19 (EN) → +64.1% tokens per character. Output: same question answered in Indonesian used ≥ 1.40× the tokens (espresso-ID hit the 450 cap in all reps, so the true ratio is at least this).
6 Each few-shot example costs ≈ 17.3 tokens for a short Text/Label pair (linear, r² = 0.9990). Budget examples deliberately: 5 shots ≈ 86 extra tokens on every single call.
7 Digit strings are token-dense: ≈ 2.0 chars/token (501 tokens per 1,000 chars) vs 4.76 for English prose. Python code averaged 2.70 ± 0.47 chars/token. Sending long numeric dumps is disproportionately expensive — prefer aggregates.
8 The tokenizer is deterministic. Every duplicated payload in the plan measured an identical count (json_min 5,130 = 5,130; pretty JSON measured three times: 5,831 / 5,831 / 5,831).

Charts (all generated from the raw data by benchmark/analyze.py):

formats scaling
density verbosity

Full tables: results/SUMMARY.md · flat data: results/summary.csv · raw responses: results/results.json

A practical playbook (backed by the numbers above)

  1. Ship tabular data as CSV/TSV, not XML or pretty JSON (−68.5% on identical data).
  2. If it must be JSON, minify it (−42% vs indent-2 on this dataset). If humans must read it, note that indent depth beyond 2 is free — go 8 if you like.
  3. Add an explicit length instruction ("answer in at most 2 sentences") when you don't need an essay: −72.9% output tokens in this run, and output tokens are the expensive ones.
  4. Treat every few-shot example as a ~17-token recurring tax and prune to the minimum that holds quality.
  5. Working in Indonesian (or similarly tokenizer-sparse languages)? Budget ~1.6× input and ≥ 1.4× output for the same content — or keep data/keys in English and only the final answer in Indonesian.
  6. Don't paste raw numeric dumps (≈ 2 chars/token); send statistics or aggregates instead.
  7. On claude.ai-mediated channels, remember the platform's own ≈ 4.2k-token envelope rides on every call — and always verify the model field of the response, not just what you requested.

What exactly was run

82 frozen API calls (plan hash 8ad81412906cd143, 100% success, 380,962 tokens measured end-to-end on 2026-08-01):

Experiment Calls Design
cal — request-envelope calibration 1 single-character probe
e1 — token density by content type 12 EN prose / ID prose / Python / digit strings, 3 samples each
e2 — serialization formats 8 one 25-record dataset × 8 formats
e2b — scaling 12 3 formats × {5, 10, 25, 50} records → linear fit
e3 — JSON indentation 4 indent ∈ {0, 2, 4, 8}
e6 — few-shot marginal cost 6 0…5 examples → linear fit
e4 — output-brevity control 27 3 questions × {none, user-brief, concise-system} × 3 reps
e5 — question language EN vs ID 12 2 questions × 2 languages × 3 reps

Input-side experiments are max_tokens: 1 probes that read usage.input_tokens (cost ≈ 1 output token each); output-side experiments are real generations reading usage.output_tokens. Full design, assumptions, and threats to validity: docs/METHODOLOGY.md.

Reproduce it

Browser (no install): open web/token-lab.html — a self-contained harness with a live instrument panel. Outside claude.ai, paste your own API key (it never leaves your browser). Inside claude.ai it runs keylessly through the platform channel (which, as this run shows, pins its own model).

CLI (any model, e.g. the originally-targeted claude-fable-5):

pip install -r requirements.txt
ANTHROPIC_API_KEY=sk-ant-...  python benchmark/run_benchmark.py --model claude-fable-5
python benchmark/analyze.py          # rebuilds SUMMARY.md, summary.csv, charts/

run_benchmark.py writes incrementally and resumes if interrupted. The plan itself is generated deterministically by benchmark/datasets.py; CI verifies the committed experiments.json matches a fresh regeneration.

Repository layout

benchmark/   datasets.py (frozen plan generator) · run_benchmark.py (CLI)
             analyze.py (stats + charts) · build_artifact.py · experiments.json
web/         token-lab.html — self-contained browser harness
results/     results.json (raw) · SUMMARY.md · summary.csv · charts/*.png
docs/        METHODOLOGY.md

Limitations, honestly

Small n (3 reps per generation cell) on a single day; six generation replies hit the 450-token cap (all in two Indonesian no-instruction cells) so those means are right-censored lower bounds; the served model was claude-sonnet-4-6, not the requested claude-fable-5; token densities depend on the specific samples (which ship verbatim in experiments.json for audit); tokenizers and channel envelopes can change over time. Numbers here are honest measurements of one configuration on one day — rerun before betting money on them.

Contributing

Reruns on other models (especially claude-fable-5 via the CLI), added experiments, or corrections are very welcome — see CONTRIBUTING.md. This is an independent community measurement project, not affiliated with or endorsed by Anthropic.

License

MIT © 2026 Alfonsus Abdi


🇮🇩 Bahasa Indonesia

Statistik efisiensi token API Claude yang benar-benar diukur — 82 panggilan API sungguhan dengan payload yang dibekukan (hash 8ad81412906cd143), data mentah lengkap, dan analisis yang bisa direproduksi. Setiap angka diambil dari field usage respons API asli; tidak ada yang diestimasi.

Model yang terukur: rencana meminta claude-fable-5, tetapi jalur artifact claude.ai yang menjalankan pengukuran ternyata melayani claude-sonnet-4-6 — tercatat di field model pada seluruh 82 respons (itulah sumber kebenarannya). Jalur ini juga menyematkan amplop ≈ 4.172 token (system prompt platform) pada setiap panggilan; seluruh tabel input sudah dikoreksi terhadap amplop terukur ini. Siapa pun bisa mengulang pengukuran pada claude-fable-5 lewat CLI dengan API key sendiri.

Temuan utama (2026-08-01)

  1. CSV/TSV format termurah untuk data tabular, XML termahal: 25 record yang sama = 618 vs 1.963 token → XML 3,18× CSV (hemat 68,5%).
  2. JSON rapi (+indentasi) 73% lebih boros daripada minified — tetapi memperdalam indentasi 2→4→8 tidak menambah token sama sekali (958 → 1.659 → 1.659 → 1.659).
  3. Biaya marginal per record: csv 24,3 · tabel markdown 33,3 · JSON rapi 66,3 token/record (r² = 1,0000).
  4. Satu kalimat instruksi "jawab maksimal 2 kalimat" memangkas output −72,9% (349,4 ± 89,3 → 94,6 ± 27,2 token); system prompt ringkas −63,4%.
  5. Bahasa Indonesia lebih boros token: input 2,90 karakter/token vs Inggris 4,76 (+64,1% token per karakter); jawaban berbahasa Indonesia memakai ≥ 1,40× output token untuk pertanyaan yang sama (dua sel ID mentok batas 450 token, jadi rasio aslinya minimal segitu).
  6. Tiap contoh few-shot ≈ 17,3 token (r² = 0,9990) — pajak berulang di setiap panggilan.
  7. Deretan angka sangat boros (≈ 2,0 karakter/token; 501 token per 1.000 karakter) — kirim agregat, bukan dump mentah.
  8. Tokenizer deterministik: semua payload duplikat terukur identik.

Tabel lengkap ada di results/SUMMARY.md (berbahasa Indonesia), metodologi rinci + keterbatasan di docs/METHODOLOGY.md.

Reproduksi

Buka web/token-lab.html di browser (di luar claude.ai isi API key sendiri — key tidak pernah meninggalkan browser), atau lewat terminal:

ANTHROPIC_API_KEY=sk-ant-...  python benchmark/run_benchmark.py --model claude-fable-5
python benchmark/analyze.py

Catatan

Proyek pengukuran independen komunitas; tidak berafiliasi dengan dan tidak disponsori Anthropic. Ukuran sampel kecil (3 ulangan per sel) dan diambil pada satu hari — ulangi pengukuran sebelum dipakai untuk keputusan biaya besar.

About

Measured token-efficiency statistics for the Claude API: 82 real calls, frozen payloads, raw data, reproducible analysis.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages