VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
-
Updated
Jul 8, 2026 - Python
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
开源 AI 视频本地化工具:自动完成 YouTube/Bilibili 视频下载、字幕识别与翻译、语音克隆配音、音轨混合和字幕压制。
Talk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.
TTS-Story is a web-based multi‑voice TTS studio for turning tagged scripts into audiobooks—featuring full speaker management, chunk review/regeneration, a job queue and library system, and local GPU or API backends including Kokoro, Chatterbox, VOX CPM, Pocket-TTS, Kitten-TTS, IndexTTS-2, QWEN3 TTS and Omnivoice engines
A clean, efficient ComfyUI custom node for VoxCPM TTS (Text-to-Speech) functionality. This implementation provides high-quality speech generation and voice cloning capabilities using the VoxCPM 1.5 model.
多模型语音合成平台 - 支持 VoxCPM2、IndexTTS 2.0、dots.tts 三引擎,声音克隆、声音设计、LoRA 微调、多角色剧本配音,9 种预置音色,中英日韩多语言界面,一键部署 | Multi-model TTS platform: VoxCPM2 + IndexTTS 2.0 + dots.tts, voice cloning, voice design, LoRA fine-tuning, multi-character dubbing. 9 preset voices, zh/en/ja/ko UI. FastAPI + HTMX, one-click Docker deploy.
DubCue is a local AI dubbing director for script-based voiceover workflows, currently built around VoxCPM2 with multi-model local TTS support in development.
Multi-engine TTS server (Qwen3-TTS + VoxCPM2): MLX & PyTorch backends, voice cloning, incremental PCM streaming, one-model-in-VRAM manager with idle eviction.
Build voice apps fast. Unified API for speech recognition & synthesis with streaming, WebSocket, and multi-engine support.
One-click Pinokio installer for VoxCPM2 with voice cloning api and prompt memory.
⚡ A talking desktop Pikachu that runs 100% on-device — MiniCPM5-1B brain + VoxCPM voice + Nemotron ears, every model ≤1B params. No cloud, works with the Wi-Fi off.
One-click Pinokio launcher for VoxCPM2 Portable. Multilingual TTS (30 languages), Voice Design, Voice Cloning, end-to-end LoRA fine-tuning. Cross-platform (Win/Linux/macOS, NVIDIA/AMD/CPU).
Standalone C++ inference project for VoxCPM models built on top of ggml with API Frontend
A local-first macOS personal radio player with SwiftUI, Apple Music, VoxCPM2 TTS, and a lightweight DJ sidecar.
100% Free & Unlimited ElevenLabs Alternative for Windows PC (Offline AI Voice Cloning & TTS) - Creative Audio AI
Local VoxCPM2 text-to-speech API server for Apple Silicon (MLX) — voice design + voice cloning, OpenAI-style endpoint
VoiceGen Oui!: Windows + WSL2 bridge for ROCm-powered VoxCPM2 voice generation on AMD GPUs. RX 7900 XTX verified; AMD GPU testers welcome.
Reverse-Turing webcam interrogation game. An AI interrogator (A.M.N.) uses webcam, microphone, pulse detection, and micro-expression analysis to determine if you're human.
Add a description, image, and links to the voxcpm topic page so that developers can more easily learn about it.
To associate your repository with the voxcpm topic, visit your repo's landing page and select "manage topics."