Skip to content
llm-measurementPublic

About

Investigate LLM token spikes and usage anomalies. Find leading users and sessions, compare windows, and check missing usage locally.

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

fleetdiff

OpenSSF Scorecard

fleetdiff is a local command-line tool that investigates LiteLLM spend exports, compares LLM usage summaries, flags unusual windows, and checks trace captures and collector configurations.

Use it to answer why LLM token usage spiked, and which users or sessions drove the change. It reads your files and prints a report: no account, upload, collector or model API key is needed for the LiteLLM path.

Commands investigate, scan, compare, inspect, diagnose
Platforms Linux and macOS, AMD64 and ARM64
Status Pre-1.0; see the changelog
License Apache-2.0

Using LiteLLM?

Export -> install -> investigate. Use the request logs you already retain to compare two periods, see which models and keys drove the change, and check whether the usage records are complete.

1. Export Request-Level Rows

Use the read-only SQL projection and your existing PostgreSQL access. The export guide covers connection settings and optional cache fields. Once your approved read-only connection is named litellm_readonly, this example exports 16 UTC days:

umask 077
export_dir="$HOME/litellm-export"
mkdir "$export_dir"
curl --fail --location --proto '=https' --tlsv1.2 \
  https://raw.githubusercontent.com/llm-measurement/fleetdiff/v0.6.0/examples/litellm-spend/export.sql \
  --output "$export_dir/export.sql"
psql 'service=litellm_readonly' -Xq -v ON_ERROR_STOP=1 \
  -v start_utc=2026-09-23T00:00:00Z -v end_utc=2026-10-09T00:00:00Z \
  -v csv_file="$export_dir/spend.csv" -f "$export_dir/export.sql" \
  >"$export_dir/private-psql.out" 2>"$export_dir/private-psql.err"

Choose dates for your own data and continue only after the export succeeds. The SQL writes CSV in one pass and excludes prompts and responses. Keep the export directory private. Existing request-level JSON and JSONL work too.

2. Install v0.6.0

Run this in Bash with curl, gh, shasum and tar available:

bash -o pipefail -c 'curl --fail --location --proto "=https" --tlsv1.2 https://raw.githubusercontent.com/llm-measurement/fleetdiff/e75518e1b62ead77a6870cfeabd3ddae944a1e0a/scripts/install.sh | sh -s -- v0.6.0 ./fleetdiff-install'
./fleetdiff-install/fleetdiff --version

The installer checks signed release provenance and checksums before extracting the binary into a new directory. Prefer to inspect the script first? Follow Install and verify.

3. Investigate

./fleetdiff-install/fleetdiff investigate \
  --litellm-spend "$export_dir/spend.csv" --group-by key

With enough history, fleetdiff compares the two most recent complete seven-day UTC periods inside the file and prints their exact boundaries and weekdays. For shorter exports or chosen periods, set --before-period and --after-period to dates or explicit UTC start/end intervals.

The report shows request count versus tokens per request by model, the leading contributors, and usage completeness. Choose --group-by team, user, end-user or session for another view; add --format json for automation. Full command and export guide.

Try One Example

Download the synthetic spend file and use the installed binary:

curl --fail --location --proto '=https' --tlsv1.2 \
  https://raw.githubusercontent.com/llm-measurement/fleetdiff/v0.6.0/examples/litellm-spend/synthetic.csv \
  --output synthetic.csv
./fleetdiff-install/fleetdiff investigate --litellm-spend synthetic.csv \
  --before-period 2026-10-07 --after-period 2026-10-08

Selected output:

Recorded tokens doubled (8,000 -> 16,000).
model-1 accounts for 100% of the net recorded increase (+8,000 tokens).
  model-1: 4,000 -> 12,000 recorded tokens; 4 -> 6 requests.
    1,000.00 -> 2,000.00 recorded tokens per request; +3,000 from request count, +5,000 from request size.
  2 leading tracked keys account for 100% of the net recorded increase.
  Zero-only records, origin unknown: 2 -> 4.
  Failed records (overlap usage categories): 1 -> 2.

Two keys explain the increase. The zero-only and failed rows point to records worth reviewing in LiteLLM. The fixture and checks make these answers reproducible.

The summary-file path also shows session concentration:

Synthetic session investigation: one of eight tracked sessions flagged for review, accounting for 90.91% of attributed tokens.

Run the sessions demo | Read its transcript.

Install

The verified installer above selects the native binary for your machine. For manual downloads, use the v0.6.0 archives and verification instructions. An administrator can verify and transfer artifacts through an approved internal mirror.

With a security-patched Go 1.26 toolchain, install the same version from its module:

go install github.com/llm-measurement/fleetdiff/cmd/fleetdiff@v0.6.0

The executable goes to GOBIN, or $(go env GOPATH)/bin when unset. See source builds and version reporting.

Limits

  • Spend reports describe the supplied logged model requests. Missing usage, ambiguous zero-only counts and failures remain visible; unsupported call types lead with an incomplete comparison. Provider records remain the source for billing.
  • Use request-level exports. The LiteLLM Usage dashboard's daily aggregates receive instructions for getting those rows. The reader accepts up to 10M rows and 16 GiB per file, with per-record limits.
  • Reports use local aliases or keyed hashes. Hashes are pseudonymous and linkable; keep source files private and share reports under your organization's policy.
  • Rankings carry deterministic bounds; distinct counts are statistical estimates. A flagged contributor identifies something to investigate, rather than its cause. Compatible summary files need aligned windows, keys and disjoint producer coverage.

Documentation

Questions or feedback: open an issue. See Contributing for checks and signed, signed-off commits. Code authors: Vijay and Codex.

About

Investigate LLM token spikes and usage anomalies. Find leading users and sessions, compare windows, and check missing usage locally.

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages