Skip to content

Design spec pitch: piper CPU TTS fallback tier (benchmark-first) - #386

Merged
tonythethompson merged 2 commits into
mainfrom
vibe/piper-cpu-tts-tier-dfc30d
Oct 6, 2026
Merged

tonythethompson merged 2 commits into
mainfrom
vibe/piper-cpu-tts-tier-dfc30d

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Revised per review: benchmark Kokoro against piper on target low-end CPUs before designing automatic fallback; Phase 0 measures Kokoro RTF, characterizes real-world espeak-ng failure cases, and quantifies the piper RTF/naturalness trade on the same hardware
  • Fallback selection is owned by the runtime planner (RoutedTtsEngine → StageRuntimePlanningRequest/RuntimePlanner), with piper ranked below Kokoro in TTS stage requirements — no engine-internal fallback logic
  • The embedded-eSpeak-NG question is a hard gate: piper only addresses Kokoro's espeak fragility if its own phonemizer and data are packaged and health-checked (same rigor as EspeakNgHealthCheck); GPL implications of embedded eSpeak-NG are a Phase 0 legal exit criterion, and per-voice model-card review with commercial_use_verified is required per manifest schema and repo governance
  • No-Python requirement: the C++ inference core is the integration target; no end-user runtime dependencies per repo policy — upstream standalone-binary feasibility is a Phase 1 blocker if unsupported
  • Removed the duplicated L1 spec file from this branch (it is now single-commit off main, containing only the piper spec)

Verification

  • Docs-only change; no build impact
  • Evidence citations verified against live code: RoutedTtsEngine planner path, EspeakNgHealthCheck, manifest schema required fields, AGENTS.md model governance

Copilot AI balanced review requested due to automatic review settings October 6, 2026 10:41
@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

How to use the Graphite Merge Queue

Add either label to this PR to merge it via the merge queue:

  • queue - adds this PR to the back of the merge queue
  • fast - for urgent changes, fast-track this PR to the front of the merge queue

You must have a Graphite account in order to use the merge queue. Sign up using this link.

An organization admin has enabled the Graphite Merge Queue in this repository.

Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@amazon-q-developer amazon-q-developer Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR introduces two comprehensive design specifications for CPU-tier inference capabilities. Both specs are well-structured with clear problem statements, goals, execution plans, risk analyses, and open questions.

Key Strengths:

  • Clear problem definition and target segment identification
  • Thoughtful architectural decisions (managed sidecar pattern, loopback-only, health checks)
  • Appropriate quality gates and benchmarking requirements before adoption
  • Comprehensive risk mitigation strategies
  • Staged execution plans with evidence-based decision gates

The specs appropriately position these as emergency fallback/CPU-only tiers without replacing existing GPU-optimized paths, and include proper readiness semantics consistent with the existing "never-fake-readiness" invariant. No blocking defects identified for merge.


You can now have the agent implement changes and create commits directly on your pull request's source branch. Simply comment with /q followed by your request in natural language to ask the agent to make changes.

@tonythethompson tonythethompson left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs-only review of the piper TTS fallback pitch. Five threads: the llama.cpp spec duplicated from #385, a wrong type name, piper's own espeak-ng dependency (which undercuts the main goal), a missing license/commercial-use section, and two design points on in-process ONNX vs sidecar and where fallback selection should live.

@@ -0,0 +1,145 @@
# Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This file is byte-identical to the one in #385 (same blob c0f9401a), and #385 already has open review threads on it (wrong model-size premise, CPU already being an allowed TextRefinement provider, the nonexistent DnnlEngineCacheProbe, routing living in the runtime planner). Carrying it here means whichever PR merges first lands it unreviewed and the other one conflicts or diverges.

Ask: drop this file from #386 and either stack #386 on #385's branch or just keep the cross-reference by path.

Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
`KokoroTtsEngine` is already the CPU-friendly default, so this is **not** a quality-tier upgrade pitch. The gaps are narrower but real:

1. **Very weak machines.** Kokoro + espeak-NG on an old/low-power CPU (or a constrained VM/mini-PC) can be too slow for practical dubbing work.
2. **espeak-ng dependency fragility on Linux.** `KokoroEspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

KokoroEspeakNgPathResolver doesn't exist. The resolver is EspeakNgPathResolver in src/Trackdub.Inference.Onnx/Kokoro/EspeakNgPathResolver.cs, next to EspeakNgHealthCheck.

Suggested change
2. **espeak-ng dependency fragility on Linux.** `KokoroEspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all.
2. **espeak-ng dependency fragility on Linux.** `EspeakNgPathResolver`/`EspeakNgHealthCheck` exist precisely because this chain breaks; when it does, the CPU user has *no* working TTS at all.

Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
## 2. Goals / Non-goals

**Goals**
- A last-resort local TTS tier that works on virtually any CPU, with no espeak-ng-style fragile native dependency chain.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This goal doesn't hold as written: piper phonemizes with espeak-ng too (piper-phonemize links libespeak-ng and needs an espeak-ng-data directory, and piper voices' phoneme_id_map is keyed on espeak IPA). So piper doesn't remove the espeak-ng chain; at best it ships its own bundled copy.

That matters for gap 2 and open question 3: the failures EspeakNgHealthCheck reports today (executable not resolvable, espeak-ng-data missing) are exactly the class a piper fallback can also hit. If the real fix for gap 2 is "ship a known-good espeak-ng", EspeakNgPathResolver already supports bundled locations, and doing that for Kokoro fixes gap 2 without a new tier.

Ask: reword this goal to say piper bundles its own espeak-ng, and split the motivation so gap 2 (packaging/bundling espeak-ng) is handled separately from gap 1 (speed on very weak CPUs), which is the part piper actually solves.

Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
└─ PiperTtsEngine (NEW — ultra-light fallback tier)
```

- **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things on the sidecar choice.

First, piper voices are plain VITS .onnx models plus a .onnx.json config. Trackdub already runs ONNX Runtime in-process and already shells out to espeak-ng for phonemes (EspeakNgPhonemizer, per ADR-0005). A PiperTtsEngine could load the voice through the existing ORT session plumbing and reuse that phonemizer process, with no new sidecar binary, no HTTP surface, and no dependency on L1's sidecar host shipping first (which the Risks table currently uses to defer this).

Second, the upstream C++ piper binary is a stdin/stdout CLI; the HTTP server is the Python package, so "minimal CLI/HTTP interface" plus "loopback-only API with health checks" implies shipping a Python runtime or writing our own wrapper.

Ask: evaluate in-process ORT + existing phonemizer as the default option and only keep the sidecar if there's a concrete reason it's needed.

Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
```

- **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean.
- **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's no licensing/commercial-use section, and this repo has precedent for needing one (ADR-0005 on espeak-ng's GPL boundary, ADR-0006 for Chatterbox commercial use).

  • The binary: piper links libespeak-ng (GPL-3.0-or-later), and active development moved from the archived rhasspy/piper to OHF-Voice/piper1-gpl, which is GPL-3.0. Per ADR-0005 that's only OK across a process boundary, so the spec should state piper (or its phonemizer) must never be loaded in-process.
  • The voices: each piper voice has its own MODEL_CARD with the training dataset's license, and they aren't uniform, so some may not be cleared for commercial use.

Ask: add a license section and make a per-voice commercial-use check a requirement for every manifest entry, the same way ADR-0006 did for Chatterbox.

Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated

- **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean.
- **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded.
- **Automatic fallback with transparency.** Default routing picks the tier by hardware and health: Kokoro fails its health check or benchmarks unacceptably slow on the detected tier → piper is selected, and the readiness panel/summary shows *why* (per the never-fake-readiness invariant — the user must see they're on the fallback voice tier).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RoutedTtsEngine doesn't pick engines by health itself. It asks IRuntimePlanner.PlanAsync for a StageRuntimePlan and then InferenceEngineAdapterSelector.SelectForPlan picks the adapter. So the Kokoro-to-piper fallback (and the "benchmarks unacceptably slow" trigger) belongs in the runtime planner for RuntimeStage.Tts, not in a new TTS router or in the health check.

This also covers the transparency requirement for free: RoutedTtsEngine.BuildSummary already copies plan.Fallback?.Detail into the stage summary as BootstrapDetail, so a planner-level fallback with a detail string like "Kokoro: espeak-ng unavailable" surfaces in summaries without new plumbing.

Ask: point this section and execution step 2 at the runtime planner and the existing plan.Fallback path.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The specs contain scope, routing, governance, and internal consistency issues that must be corrected before approval.

Review effort: Balanced
Findings: 11 Low severity

Open (11)
What changed in this PR

Adds design proposals for low-resource CPU fallbacks across TTS, text refinement, translation, and ASR.

Changes:

  • Proposes a Piper TTS fallback tier.
  • Proposes llama.cpp and whisper.cpp CPU tiers.
  • Defines routing, readiness, benchmarking, and distribution plans.
File Description
docs/​specs/​design-piper-cpu-tts-tier.md Specifies the Piper fallback architecture and rollout.
docs/​specs/​design-llamacpp-cpu-text-tier.md Specifies CPU-side text, translation, and ASR runtimes.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@@ -0,0 +1,145 @@
# Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR
# Design Spec — L1: llama.cpp GGUF CPU Tier — Text Refinement, Translation, ASR

**Pitch status:** proposed — seeking product sign-off before any code.
**One-liner:** make local text refinement, translation, and ASR real for GPU/NPU-less users by adding llama.cpp/GGUF engines behind the existing stage routers, with hardware-tier-aware default routing. One sidecar runtime, three stages.
| `QwenTextRefinementEngine` | ONNX Runtime GenAI; accelerated via TensorRT-RTX / NPU EPs (`tools/olive/Export-QwenTrtRtxGenAi.ps1`, VitisAI/QNN/OpenVINO recipes) | Mid-range RTX GPU or supported NPU |
| `GeminiCloudTextRefinementEngine` | Cloud API | Network + egress consent |

`RoutedTextRefinementEngine` (`src/Trackdub.Inference.Onnx/QwenTextRefinement/RoutedTextRefinementEngine.cs`) picks the Qwen local engine by default and the Gemini engine only on explicit alias. **There is no local option for a CPU-only or low-VRAM machine.** On that hardware, ORT GenAI CPU execution of the current 7–9B class export is slow enough that users silently fall back to cloud — or abandon local refinement entirely. GPU/NPU-less users are an explicit target segment, so this is a coverage gap, not an edge case.

**Non-goals**
- Replacing ONNX/TRT-RTX on hardware that already runs it well.
- llama.cpp for any stage other than text refinement (ASR/TTS stay as-is).

**Model manifest.** Extend the existing model catalog/readiness flow (`EnsureTextRefinementModelAvailableAsync` in the stage coordinator) with GGUF entries: alias, GGUF URL + size + SHA-256, recommended quant per hardware tier, prompt-template metadata. A curated set of one model family (same Qwen family as the current engine, for output-guard comparability) in Q4_K_M and Q5_K_M is enough for v1.

**Hardware tier detection** reuses existing capability probing (`Trackdub.Inference.Onnx/Runtime` EP probes, `DnnlEngineCacheProbe`-style pattern) to classify:

## 6. L1-T: Translation — the same sidecar, one more engine *(highest incremental ROI)*

The high-quality translation route today is `phi-genai` (`PhiGenAiTranslationEngine`, ORT GenAI LLM — the same hardware floor as text refinement). CPU-only users get routed to the small encoder routes (`madlad`, `opus-mt`) via `TranslationLanguageRouter` — a real quality cliff, and the most user-visible one in the pipeline.
3. **L1-R quality gate + benchmark:** `QwenRefinementOutputGuard` suite at Q4_K_M/Q5_K_M vs INT4; DubBench tokens/sec + TTFT on reference CPU machine.
4. **L1-T:** `LlamaCppTranslationEngine` adapter on the now-proven sidecar/manifest foundation; translation-quality suite; `TranslationLanguageRouter` tier wiring.
5. **L1-A benchmark (evidence gate):** CPU ASR DubBench comparison; ship whisper.cpp tier only if the measured gap clears the threshold.
6. **Distribution:** sidecar acquisition story (bundled vs on-demand download) — decided in review; on-demand matches the existing model-download UX.
Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
```

- **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean.
- **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded.
Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated

- **Sidecar, not in-process.** Same managed-sidecar pattern as L1: piper ships a small standalone binary with a minimal CLI/HTTP interface; run it loopback-only as a child process with health checks. Keeps the native dependency out of the .NET process and keeps the `Trackdub.Inference.*` dependency direction clean.
- **Model manifest.** Piper voice models are small single files (tens of MB) — a curated set (one voice per supported language for v1) rides the same manifest/readiness machinery built for L1: sidecar present → voice downloaded/hash-verified → process healthy → stage ran → succeeded.
- **Automatic fallback with transparency.** Default routing picks the tier by hardware and health: Kokoro fails its health check or benchmarks unacceptably slow on the detected tier → piper is selected, and the readiness panel/summary shows *why* (per the never-fake-readiness invariant — the user must see they're on the fallback voice tier).
Comment thread docs/specs/design-piper-cpu-tts-tier.md Outdated
## 4. Execution plan

1. **Spike (acceptance slice):** `PiperTtsEngine` implementing the existing TTS engine contract + sidecar process manager (reusing the L1 shared sidecar host code) + one voice model manifest entry; manual alias selection only, behind a feature flag.
2. **Health-driven fallback:** extend the existing TTS readiness/health checks (modeled on `EspeakNgHealthCheck`) so Kokoro degradation/failure automatically selects piper, with the fallback surfaced in stage summaries and the readiness panel.
…ense gates

Co-authored-by: tonythethompson <tonythethompson@users.noreply.github.com>
@tonythethompson
tonythethompson force-pushed the vibe/piper-cpu-tts-tier-dfc30d branch from 13def8f to 7e42d77 Compare October 6, 2026 14:57
@tonythethompson tonythethompson changed the title Design spec pitch: piper ultra-light CPU TTS fallback tier Design spec pitch: piper CPU TTS fallback tier (benchmark-first) Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Phase 0 benchmark results — Kokoro vs piper (constrained CPU)

Environment: Linux sandbox, 5 vCPU Intel Xeon 6985P-C @ 2.30GHz (no GPU), 5.9 GB RAM. Caveat: this is a server-class constrained-core environment, not a true low-end consumer laptop (no thermal throttling, laptop single-core boost behavior) — it bounds the relative engine comparison but should be re-run on the actual target hardware profile before final sign-off.

Setups: Kokoro-82M ONNX fp32 export (onnx-community/Kokoro-82M-v1.0-ONNX, model.onnx, voice af_heart) via kokoro-onnx 0.6.1 + ORT 1.30 CPU EP — the ONNX path mirrors KokoroTtsEngine's architecture. Piper en_US-lessac-medium (63 MB ONNX, embedded espeak-ng-data 19 MB). 3 warmup + 3 repeats, best-of; 5-sentence dubbing-style corpus (short to long).

Metric piper (lessac-medium) Kokoro (fp32 ONNX)
Model load 0.78 s 0.81 s
Total synth (23 s audio, 5 sentences) 0.65 s 4.64 s
Aggregate RTF (wall / audio) 0.028 0.208
Short-segment RTF ("Watch this.") 0.028 0.338
Long-sentence RTF 0.026 0.187
Peak RSS (long sentence, warm) 327 MB 807 MB
On-disk footprint 63 MB model + 19 MB espeak data 325 MB model (fp32; quantized exports exist down to ~80 MB q4)

Findings:

  1. Kokoro RTF 0.19-0.34 on constrained cores — comfortably under 1. It already synthesizes ~5x faster than real time. Per the spec's decision matrix, the speed concern does not justify a piper tier on this hardware class.
  2. Piper is ~7x faster and ~2.5x lighter in RAM, as expected for an ultra-light tier — but that headroom is not currently needed for playback-rate synthesis. The remaining case is only extreme hardware (very weak single/dual-core, under 2 GB free RAM) or heavy batch/offline throughput, which this environment cannot represent.
  3. espeak-ng dependency: both engines sit on eSpeak-NG for phonemization in this benchmark. Kokoro (kokoro-onnx) phonemizes via phonemizer/espeakng-loader (pip-shipped libespeak-ng.so + espeak-ng-data); piper ships the same machinery embedded. So piper does not de-risk the espeak chain by substitution — it also depends on it, just with self-contained packaging. This confirms the spec's hard gate: the fragility question is a packaging/health-check problem for either engine, not a reason to switch engines. Trackdub's native EspeakNgPathResolver/EspeakNgHealthCheck already owns this correctly for Kokoro.
  4. Quality note (not measured here): piper at RTF 0.028 is audibly lower-naturalness than Kokoro. The trade is real but currently one-sided — nothing in the speed data demands taking it.

Conclusion against the spec's decision matrix (§3): on this hardware class the evidence lands on "Kokoro RTF fine, fragility = packaging question" -> close the spec / no piper tier; prefer hardening Kokoro's espeak path (bundle espeak-ng-data, keep the health check) as the spec already anticipates. Re-validate on a true low-end laptop profile if that segment still feels underserved — a Raspberry Pi-class or 2-core/4 GB machine would be the right target.

Benchmark harness + raw JSON can be dropped into DubBench for the reference-profile re-run.

…trained CPU

Co-authored-by: tonythethompson <tonythethompson@users.noreply.github.com>
@tonythethompson
tonythethompson merged commit 0472898 into main Oct 6, 2026
20 of 21 checks passed
@tonythethompson
tonythethompson deleted the vibe/piper-cpu-tts-tier-dfc30d branch October 6, 2026 17:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants