Repository navigation
Add tts-bench command: Kokoro TTS benchmarking for DubBench - #388
Conversation
How to use the Graphite Merge QueueAdd either label to this PR to merge it via the merge queue:
You must have a Graphite account in order to use the merge queue. Sign up using this link. An organization admin has enabled the Graphite Merge Queue in this repository. Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue. |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
This PR successfully integrates a TTS benchmarking command (tts-bench) into the Trackdub.Benchmarks tooling. The implementation follows established patterns from SeparationEvalRunner, includes comprehensive test coverage (18 tests), and properly routes through the real product path via HeadlessDubbingHost and KokoroTtsEngine.
Key strengths:
- Clean separation of concerns with dedicated records for jobs, options, and results
- Robust error handling with graceful failure continuation
- Proper resource management with
WorkingSetPeakMonitorand disposal patterns - Comprehensive test coverage mirroring existing benchmark test patterns
- JSONL streaming output for incremental result capture
The code is well-tested (18/18 tests passing), follows the repository's conventions, and integrates cleanly with the existing benchmark infrastructure. No blocking issues identified.
You can now have the agent implement changes and create commits directly on your pull request's source branch. Simply comment with /q followed by your request in natural language to ask the agent to make changes.
tonythethompson
left a comment
There was a problem hiding this comment.
Two things that affect whether the numbers mean what the PR says they mean; inline below.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Model pinning, alias handling, per-job validation, and provider usage text can currently produce incorrect or unusable benchmarks.
Review effort: Balanced
Findings: 3
Open (3)
What changed in this PR
Adds a Kokoro TTS benchmarking command for collecting timing, RTF, memory, and provider evidence.
Changes:
- Adds JSONL-driven TTS benchmark execution and reporting.
- Registers
tts-benchand adds CLI usage. - Adds tests and the related Piper benchmark-first design spec.
| File | Description |
|---|---|
src/Trackdub.Benchmarks/TtsEvalRunner.cs |
Implements parsing, execution, and JSONL results. |
src/Trackdub.Benchmarks/Program.cs |
Dispatches the new command. |
src/Trackdub.Benchmarks/BenchmarkConsole.cs |
Documents command usage. |
tests/Trackdub.Benchmarks.Tests/TtsEvalRunnerTests.cs |
Tests parsing and execution behavior. |
docs/specs/design-piper-cpu-tts-tier.md |
Documents the benchmark-first Piper proposal. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
OpenVINO EP vs CPU EP smoke results (Intel CPU sandbox)Environment: same constrained-CPU sandbox as the TTS benchmark — Intel Xeon 6985P-C @ 2.30 GHz, no GPU/NPU. Stack: 1. EP discovery & session creation
2. whisper-base encoder (ASR-class, 80×3000 features ≈ 30 s audio), warm, best-of-5
3. Kokoro-82M (all variants: fp32, fp16, q8f16) — OpenVINO EP cannot load the graphAll three exports fail at session creation: The Kokoro graph's STFT output has dynamic rank, which the OpenVINO CPU plugin rejects outright (ORT CPU EP runs the same file fine). So for the current Kokoro TTS artifact, the OpenVINO EP is not a drop-in alternative without a graph fix (pin the STFT output rank/shape in the export). 4. Read-through to the CPU-tier spec (#385)
Raw results JSON and harness available in the benchmark working notes; can be adapted into |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
PR Summary by QodoAdd Kokoro TTS benchmark command and benchmark-first Piper spec
AI Description
Diagram
High-Level Assessment
Files changed (5)
|
Code Review by Qodo
1.
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6a641db799
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Qodo Fixer🍒 Ready to be cherry-picked — ✅ Merged (0) · ☑ Fixed (1) 🔗 Fix PR: #389 This fix PR was closed automatically. Its branch is preserved so you can cherry pick the changes into the original PR. Prompt for coding agent Process — 1 fixed
|
There was a problem hiding this comment.
All reported issues were addressed across 5 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
|
/oc fix |
|
Review posted successfully. Here is the summary: Review submitted — 5 inline findings across 2 files. Overall: The
Out of diff: No out-of-diff findings. All 5 findings are anchored inline on the PR diff. |
…el alias, warmup peak, per-job validation
|
The new commit isn't pushed yet (the action pushes automatically); the last-pushed commit SummaryConflicts — None. Merged latest Fixed (all 21 review threads resolved):
Tests — Updated sampler sequence for the warmup monitor; added Not fixed (resolved as not valid) — CI — No failures to fix; prior head commit was fully green, and the current merge + changes build and test clean locally. Checks on the new commit will run on push. |
|
/oc fix merge conflicts and failing ci tests |
|
All conflicts resolved, CI fix verified, working tree clean. Final state check done. SummaryConflicts — resolved. Merged the latest
CI — fixed. Verification:
Threads — all 28 review threads are resolved; no new unresolved feedback to triage (last-pass summary comment is the only timeline comment besides the trigger). Checks on the new commit ( |

Summary
tts-benchcommand toTrackdub.Benchmarks, adapting the Kokoro-vs-piper Phase 0 harness (evidence in Design spec pitch: piper CPU TTS fallback tier (benchmark-first) #386) into the repo's benchmark tooling for the reference-profile re-runTtsEvalRunnermodeled onSeparationEvalRunner: JSONL job list (id,text,language_code,voice_id, optionalspeed/warmup_runs/repeat_runs), runs the shippedKokoroTtsEngineviaHeadlessDubbingHost, warmup + best-of-N repeats, records wall ms (best + mean), audio seconds, RTF, working set before/peak, and provider selection per job as JSONLInferenceRequestOptionswith model alias / provider pins), so results reflect the real product path, not a synthetic loop; provider pinning via the standard--providertokensProgramcommand dispatch andBenchmarkConsoleusage; 18 new unit tests mirroringSeparationEvalRunnerTestscoverage (parse, validation, dispatch, per-job result shape, failure continuation, cancellation)Usage
Verification
dotnet build src/Trackdub.Benchmarks— 0 errorsdotnet test tests/Trackdub.Benchmarks.Tests --filter FullyQualifiedName~TtsEvalRunnerTests— 18/18 passedResourceTelemetryOptionsTestsculture tests (reproduced on clean tree without this change — sandbox lacks ICU libraries, forcing invariant globalization; not related to this PR)