Skip to content
#

tokens-per-second

Here are 43 public repositories matching this topic...

Reproducible on-device LLM benchmarks for Apple Silicon (iPhone 17 Pro, M4 Max): Apple Core AI, MLX, llama.cpp, LiteRT-LM and Core ML on the same model and harness, every number with its quantization and capture session; hybrid Mamba-2 models (Nemotron-3 Nano, Granite-4.0-H, Falcon-H1) included.

  • Updated Oct 2, 2026
  • Python

Manifest V3 Chrome Extension that tracks real-time AI response time, token count & speed (tok/s) across 17+ web AI platforms — with a Side Panel dashboard, session history, and CSV/JSON export.

  • Updated Jul 29, 2026
  • JavaScript

Add this topic to your repo

To associate your repository with the tokens-per-second topic, visit your repo's landing page and select "manage topics."

Learn more