Repository navigation
Design spec pitch: piper CPU TTS fallback tier (benchmark-first) - #386
Conversation
How to use the Graphite Merge QueueAdd either label to this PR to merge it via the merge queue:
You must have a Graphite account in order to use the merge queue. Sign up using this link. An organization admin has enabled the Graphite Merge Queue in this repository. Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue. |
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
This PR introduces two comprehensive design specifications for CPU-tier inference capabilities. Both specs are well-structured with clear problem statements, goals, execution plans, risk analyses, and open questions.
Key Strengths:
- Clear problem definition and target segment identification
- Thoughtful architectural decisions (managed sidecar pattern, loopback-only, health checks)
- Appropriate quality gates and benchmarking requirements before adoption
- Comprehensive risk mitigation strategies
- Staged execution plans with evidence-based decision gates
The specs appropriately position these as emergency fallback/CPU-only tiers without replacing existing GPU-optimized paths, and include proper readiness semantics consistent with the existing "never-fake-readiness" invariant. No blocking defects identified for merge.
You can now have the agent implement changes and create commits directly on your pull request's source branch. Simply comment with /q followed by your request in natural language to ask the agent to make changes.
tonythethompson
left a comment
There was a problem hiding this comment.
Docs-only review of the piper TTS fallback pitch. Five threads: the llama.cpp spec duplicated from #385, a wrong type name, piper's own espeak-ng dependency (which undercuts the main goal), a missing license/commercial-use section, and two design points on in-process ONNX vs sidecar and where fallback selection should live.
| @@ -0,0 +1,145 @@ | |||
| # Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR | |||
There was a problem hiding this comment.
This file is byte-identical to the one in #385 (same blob c0f9401a), and #385 already has open review threads on it (wrong model-size premise, CPU already being an allowed TextRefinement provider, the nonexistent DnnlEngineCacheProbe, routing living in the runtime planner). Carrying it here means whichever PR merges first lands it unreviewed and the other one conflicts or diverges.
Ask: drop this file from #386 and either stack #386 on #385's branch or just keep the cross-reference by path.
| `KokoroTtsEngine` is already the CPU-friendly default, so this is **not** a quality-tier upgrade pitch. The gaps are narrower but real: | ||
|
|
||
| 1. **Very weak machines.** Kokoro + espeak-NG on an old/low-power CPU (or a constrained VM/mini-PC) can be too slow for practical dubbing work. | ||
| 2. **espeak-ng dependency fragility on Linux.** `KokoroEspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all. |
There was a problem hiding this comment.
KokoroEspeakNgPathResolver doesn't exist. The resolver is EspeakNgPathResolver in src/Trackdub.Inference.Onnx/Kokoro/EspeakNgPathResolver.cs, next to EspeakNgHealthCheck.
| 2. **espeak-ng dependency fragility on Linux.** `KokoroEspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all. | |
| 2. **espeak-ng dependency fragility on Linux.** `EspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all. |
| ## 2. Goals / Non-goals | ||
|
|
||
| **Goals** | ||
| - A last-resort local TTS tier that works on virtually any CPU, with no espeak-ng-style fragile native dependency chain. |
There was a problem hiding this comment.
This goal doesn't hold as written: piper phonemizes with espeak-ng too (piper-phonemize links libespeak-ng and needs an espeak-ng-data directory, and piper voices' phoneme_id_map is keyed on espeak IPA). So piper doesn't remove the espeak-ng chain; at best it ships its own bundled copy.
That matters for gap 2 and open question 3: the failures EspeakNgHealthCheck reports today (executable not resolvable, espeak-ng-data missing) are exactly the class a piper fallback can also hit. If the real fix for gap 2 is "ship a known-good espeak-ng", EspeakNgPathResolver already supports bundled locations, and doing that for Kokoro fixes gap 2 without a new tier.
Ask: reword this goal to say piper bundles its own espeak-ng, and split the motivation so gap 2 (packaging/bundling espeak-ng) is handled separately from gap 1 (speed on very weak CPUs), which is the part piper actually solves.
| └─ PiperTtsEngine (NEW — ultra-light fallback tier) | ||
| ``` | ||
|
|
||
| - **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean. |
There was a problem hiding this comment.
Two things on the sidecar choice.
First, piper voices are plain VITS .onnx models plus a .onnx.json config. Trackdub already runs ONNX Runtime in-process and already shells out to espeak-ng for phonemes (EspeakNgPhonemizer, per ADR-0005). A PiperTtsEngine could load the voice through the existing ORT session plumbing and reuse that phonemizer process, with no new sidecar binary, no HTTP surface, and no dependency on L1's sidecar host shipping first (which the Risks table currently uses to defer this).
Second, the upstream C++ piper binary is a stdin/stdout CLI; the HTTP server is the Python package, so "minimal CLI/HTTP interface" plus "loopback-only API with health checks" implies shipping a Python runtime or writing our own wrapper.
Ask: evaluate in-process ORT + existing phonemizer as the default option and only keep the sidecar if there's a concrete reason it's needed.
| ``` | ||
|
|
||
| - **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean. | ||
| - **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded. |
There was a problem hiding this comment.
There's no licensing/commercial-use section, and this repo has precedent for needing one (ADR-0005 on espeak-ng's GPL boundary, ADR-0006 for Chatterbox commercial use).
- The binary: piper links libespeak-ng (GPL-3.0-or-later), and active development moved from the archived
rhasspy/pipertoOHF-Voice/piper1-gpl, which is GPL-3.0. Per ADR-0005 that's only OK across a process boundary, so the spec should state piper (or its phonemizer) must never be loaded in-process. - The voices: each piper voice has its own MODEL_CARD with the training dataset's license, and they aren't uniform, so some may not be cleared for commercial use.
Ask: add a license section and make a per-voice commercial-use check a requirement for every manifest entry, the same way ADR-0006 did for Chatterbox.
|
|
||
| - **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean. | ||
| - **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded. | ||
| - **Automatic fallback with transparency.** Default routing picks the tier by hardware and health: Kokoro fails its health check or benchmarks unacceptably slow on the detected tier → piper is selected, and the readiness panel/summary shows *why* (per the never-fake-readiness invariant — the user must see they're on the fallback voice tier). |
There was a problem hiding this comment.
RoutedTtsEngine doesn't pick engines by health itself. It asks IRuntimePlanner.PlanAsync for a StageRuntimePlan and then InferenceEngineAdapterSelector.SelectForPlan picks the adapter. So the Kokoro-to-piper fallback (and the "benchmarks unacceptably slow" trigger) belongs in the runtime planner for RuntimeStage.Tts, not in a new TTS router or in the health check.
This also covers the transparency requirement for free: RoutedTtsEngine.BuildSummary already copies plan.Fallback?.Detail into the stage summary as BootstrapDetail, so a planner-level fallback with a detail string like "Kokoro: espeak-ng unavailable" surfaces in summaries without new plumbing.
Ask: point this section and execution step 2 at the runtime planner and the existing plan.Fallback path.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The specs contain scope, routing, governance, and internal consistency issues that must be corrected before approval.
Review effort: Balanced
Findings: 11
Open (11)
Clarify or split the proposal's expanded scope · New Correctly describe the two-runtime, three-stage architecture · New Correct inaccurate claims about local refinement routing · New Align non-goals with the proposed translation engine · New Reference the existing DnnlReadinessProbe · New Route hardware selection through IRuntimePlanner · New Accurately describe and justify translation route changes · New Add governance gates for sidecars and bundled models · New Require licensing and attribution review for bundled voices · New Keep the core model conditional pending trigger validation · New Integrate Kokoro health into planner-based routing · New
What changed in this PR
Adds design proposals for low-resource CPU fallbacks across TTS, text refinement, translation, and ASR.
Changes:
- Proposes a Piper TTS fallback tier.
- Proposes llama.cpp and whisper.cpp CPU tiers.
- Defines routing, readiness, benchmarking, and distribution plans.
| File | Description |
|---|---|
docs/specs/design-piper-cpu-tts-tier.md |
Specifies the Piper fallback architecture and rollout. |
docs/specs/design-llamacpp-cpu-text-tier.md |
Specifies CPU-side text, translation, and ASR runtimes. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| @@ -0,0 +1,145 @@ | |||
| # Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR | |||
| # Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR | ||
|
|
||
| **Pitch status:** proposed — seeking product sign-off before any code. | ||
| **One-liner:** make local text refinement, translation, and ASR real for GPU/NPU-less users by adding llama.cpp/GGUF engines behind the existing stage routers, with hardware-tier-aware default routing. One sidecar runtime, three stages. |
| | `QwenTextRefinementEngine` | ONNX Runtime GenAI; accelerated via TensorRT-RTX / NPU EPs (`tools/olive/Export-QwenTrtRtxGenAi.ps1`, VitisAI/QNN/OpenVINO recipes) | Mid-range RTX GPU or supported NPU | | ||
| | `GeminiCloudTextRefinementEngine` | Cloud API | Network + egress consent | | ||
|
|
||
| `RoutedTextRefinementEngine` (`src/Trackdub.Inference.Onnx/QwenTextRefinement/RoutedTextRefinementEngine.cs`) picks the Qwen local engine by default and the Gemini engine only on explicit alias. **There is no local option for a CPU-only or low-VRAM machine.** On that hardware, ORT GenAI CPU execution of the current 7–9B class export is slow enough that users silently fall back to cloud — or abandon local refinement entirely. GPU/NPU-less users are an explicit target segment, so this is a coverage gap, not an edge case. |
|
|
||
| **Non-goals** | ||
| - Replacing ONNX/TRT-RTX on hardware that already runs it well. | ||
| - llama.cpp for any stage other than text refinement (ASR/TTS stay as-is). |
|
|
||
| **Model manifest.** Extend the existing model catalog/readiness flow (`EnsureTextRefinementModelAvailableAsync` in the stage coordinator) with GGUF entries: alias, GGUF URL + size + SHA-256, recommended quant per hardware tier, prompt-template metadata. A curated set of one model family (same Qwen family as the current engine, for output-guard comparability) in Q4_K_M and Q5_K_M is enough for v1. | ||
|
|
||
| **Hardware tier detection** reuses existing capability probing (`Trackdub.Inference.Onnx/Runtime` EP probes, `DnnlEngineCacheProbe`-style pattern) to classify: |
|
|
||
| ## 6. L1-T: Translation — the same sidecar, one more engine *(highest incremental ROI)* | ||
|
|
||
| The high-quality translation route today is `phi-genai` (`PhiGenAiTranslationEngine`, ORT GenAI LLM — the same hardware floor as text refinement). CPU-only users get routed to the small encoder routes (`madlad`, `opus-mt`) via `TranslationLanguageRouter` — a real quality cliff, and the most user-visible one in the pipeline. |
| 3. **L1-R quality gate + benchmark:** `QwenRefinementOutputGuard` suite at Q4_K_M/Q5_K_M vs INT4; DubBench tokens/sec + TTFT on reference CPU machine. | ||
| 4. **L1-T:** `LlamaCppTranslationEngine` adapter on the now-proven sidecar/manifest foundation; translation-quality suite; `TranslationLanguageRouter` tier wiring. | ||
| 5. **L1-A benchmark (evidence gate):** CPU ASR DubBench comparison; ship whisper.cpp tier only if the measured gap clears the threshold. | ||
| 6. **Distribution:** sidecar acquisition story (bundled vs on-demand download) — decided in review; on-demand matches the existing model-download UX. |
| ``` | ||
|
|
||
| - **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean. | ||
| - **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded. |
|
|
||
| - **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean. | ||
| - **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded. | ||
| - **Automatic fallback with transparency.** Default routing picks the tier by hardware and health: Kokoro fails its health check or benchmarks unacceptably slow on the detected tier → piper is selected, and the readiness panel/summary shows *why* (per the never-fake-readiness invariant — the user must see they're on the fallback voice tier). |
| ## 4. Execution plan | ||
|
|
||
| 1. **Spike (acceptance slice):** `PiperTtsEngine` implementing the existing TTS engine contract + sidecar process manager (reusing the L1 shared sidecar host code) + one voice model manifest entry; manual alias selection only, behind a feature flag. | ||
| 2. **Health-driven fallback:** extend the existing TTS readiness/health checks (modeled on `EspeakNgHealthCheck`) so Kokoro degradation/failure automatically selects piper, with the fallback surfaced in stage summaries and the readiness panel. |
…ense gates Co-authored-by: tonythethompson <tonythethompson@users.noreply.github.com>
13def8f to
7e42d77
Compare
Phase 0 benchmark results — Kokoro vs piper (constrained CPU)Environment: Linux sandbox, 5 vCPU Intel Xeon 6985P-C @ 2.30GHz (no GPU), 5.9 GB RAM. Caveat: this is a server-class constrained-core environment, not a true low-end consumer laptop (no thermal throttling, laptop single-core boost behavior) — it bounds the relative engine comparison but should be re-run on the actual target hardware profile before final sign-off. Setups: Kokoro-82M ONNX fp32 export (onnx-community/Kokoro-82M-v1.0-ONNX, model.onnx, voice af_heart) via kokoro-onnx 0.6.1 + ORT 1.30 CPU EP — the ONNX path mirrors KokoroTtsEngine's architecture. Piper en_US-lessac-medium (63 MB ONNX, embedded espeak-ng-data 19 MB). 3 warmup + 3 repeats, best-of; 5-sentence dubbing-style corpus (short to long).
Findings:
Conclusion against the spec's decision matrix (§3): on this hardware class the evidence lands on "Kokoro RTF fine, fragility = packaging question" -> close the spec / no piper tier; prefer hardening Kokoro's espeak path (bundle espeak-ng-data, keep the health check) as the spec already anticipates. Re-validate on a true low-end laptop profile if that segment still feels underserved — a Raspberry Pi-class or 2-core/4 GB machine would be the right target. Benchmark harness + raw JSON can be dropped into DubBench for the reference-profile re-run. |
…trained CPU Co-authored-by: tonythethompson <tonythethompson@users.noreply.github.com>

Summary
RoutedTtsEngine→StageRuntimePlanningRequest/RuntimePlanner), with piper ranked below Kokoro in TTS stage requirements — no engine-internal fallback logicEspeakNgHealthCheck); GPL implications of embedded eSpeak-NG are a Phase 0 legal exit criterion, and per-voice model-card review withcommercial_use_verifiedis required per manifest schema and repo governancemain, containing only the piper spec)Verification
RoutedTtsEngineplanner path,EspeakNgHealthCheck, manifest schema required fields, AGENTS.md model governance