Repository navigation
feat: add ModelBest VoxCPM TTS plugin #658
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
lottshin
wants to merge
3
commits into
GetStream:main
Choose a base branch
from
lottshin:feat/modelbest-voxcpm-tts
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,62 @@ | ||
| # VoxCPM Plugin | ||
|
|
||
| This package integrates the hosted [ModelBest VoxCPM](https://platform.modelbest.cn/console/docs/api/audio) Text-to-Speech API with Vision Agents. It streams audio as it is synthesized and supports optional voice cloning without requiring a local GPU. | ||
|
|
||
| ## Installation | ||
|
|
||
| ```bash | ||
| uv add "vision-agents[voxcpm]" | ||
| # or directly | ||
| uv add vision-agents-plugins-voxcpm | ||
| ``` | ||
|
|
||
| ## Usage | ||
|
|
||
| Create a ModelBest API key and select a model with the `speech_synthesis` capability, then set: | ||
|
|
||
| ```bash | ||
| export MODELBEST_API_KEY="your-api-key" | ||
| export MODELBEST_VOXCPM_MODEL_ID="your-model-id" | ||
| ``` | ||
|
|
||
| ```python | ||
| from vision_agents.plugins import voxcpm | ||
|
|
||
| tts = voxcpm.TTS() | ||
|
|
||
| try: | ||
| async for chunk in tts.send_iter("Hello from VoxCPM!"): | ||
| if chunk.data: | ||
| print(chunk.data.duration_ms) | ||
| finally: | ||
| await tts.close() | ||
| ``` | ||
|
|
||
| ## Voice cloning | ||
|
|
||
| `ref_audio` sets speaker identity. `prompt_audio` and its exact `prompt_text` can additionally preserve the reference delivery, including pacing, emotion, and pronunciation. | ||
|
|
||
| ```python | ||
| tts = voxcpm.TTS( | ||
| ref_audio="speaker.wav", | ||
| prompt_audio="delivery.wav", | ||
| prompt_text="The exact words spoken in delivery.wav.", | ||
| ) | ||
| ``` | ||
|
|
||
| Reference files must be valid, uncompressed WAV files no larger than 5 MiB. Audio supplied for cloning is sent to ModelBest; only use recordings you have permission to process and clone. | ||
|
|
||
| For a minimal live API check, see [`example/`](example/README.md). It writes the streamed response to a WAV file and can optionally test reference-audio cloning. | ||
|
|
||
| ## Configuration | ||
|
|
||
| | Parameter | Default | Description | | ||
| |---|---|---| | ||
| | `api_key` | `MODELBEST_API_KEY` | ModelBest API key. | | ||
| | `model` | `MODELBEST_VOXCPM_MODEL_ID` | Model ID with `speech_synthesis` capability. | | ||
| | `voice` | `"default"` | Protocol voice value; cloned identity comes from `ref_audio`. | | ||
| | `base_url` | `https://api.modelbest.cn/v1` | ModelBest API base URL. | | ||
| | `ref_audio` | `None` | WAV bytes or path used for speaker identity. | | ||
| | `prompt_audio` | `None` | WAV bytes or path used for delivery cloning. | | ||
| | `prompt_text` | `None` | Exact transcript paired with `prompt_audio`. | | ||
| | `request_timeout` | `120.0` | Maximum seconds without response data. | |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,5 @@ | ||
| MODELBEST_API_KEY=your-modelbest-api-key | ||
| MODELBEST_VOXCPM_MODEL_ID=your-speech-synthesis-model-id | ||
| MODELBEST_VOXCPM_BASE_URL=https://api.modelbest.cn/v1 | ||
| # Optional: path to a PCM16 WAV for speaker cloning. | ||
| VOXCPM_REFERENCE_AUDIO= |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,21 @@ | ||
| # ModelBest VoxCPM Smoke Test | ||
|
|
||
| This example makes one real request to the hosted ModelBest VoxCPM API and writes the returned speech to a WAV file. It is the smallest way to verify API credentials, model access, streamed audio parsing, and optional speaker cloning before adding VoxCPM to a full Agent. | ||
|
|
||
| ## Setup | ||
|
|
||
| ```bash | ||
| cd plugins/voxcpm/example | ||
| cp .env.example .env | ||
| ``` | ||
|
|
||
| Fill in `MODELBEST_API_KEY` and a ModelBest model ID with the `speech_synthesis` capability. Set `VOXCPM_REFERENCE_AUDIO` to a PCM16 WAV path to test speaker cloning. | ||
|
|
||
| ## Run | ||
|
|
||
| ```bash | ||
| uv run voxcpm_smoke.py | ||
| uv run voxcpm_smoke.py "你好,这是 VoxCPM 的测试语音。" | ||
| ``` | ||
|
|
||
| The output is written to `voxcpm_smoke.wav`. The command prints the number of streamed chunks, sample rate, first-chunk latency, and total duration. |
Empty file.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,13 @@ | ||
| [project] | ||
| name = "voxcpm-tts-example" | ||
| version = "0.0.0" | ||
| requires-python = ">=3.10" | ||
|
|
||
| dependencies = [ | ||
| "python-dotenv>=1.0", | ||
| "vision-agents-plugins-voxcpm", | ||
| ] | ||
|
|
||
| [tool.uv.sources] | ||
| "vision-agents-plugins-voxcpm" = { path = "..", editable = true } | ||
| "vision-agents" = { path = "../../../agents-core", editable = true } |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,67 @@ | ||
| """Make one hosted VoxCPM request and write the streamed audio to WAV.""" | ||
|
|
||
| import asyncio | ||
| import os | ||
| import sys | ||
| import time | ||
| import wave | ||
|
|
||
| from dotenv import load_dotenv | ||
| from getstream.video.rtc import PcmData | ||
| from vision_agents.plugins import voxcpm | ||
|
|
||
| load_dotenv() | ||
|
|
||
| OUTPUT_PATH: str = "voxcpm_smoke.wav" | ||
|
|
||
|
|
||
| async def main() -> None: | ||
| text = ( | ||
| sys.argv[1] | ||
| if len(sys.argv) > 1 | ||
| else "Hello from the Vision Agents VoxCPM plugin." | ||
| ) | ||
| reference_audio = os.getenv("VOXCPM_REFERENCE_AUDIO") or None | ||
| instance = voxcpm.TTS( | ||
| api_key=os.getenv("MODELBEST_API_KEY"), | ||
| model=os.getenv("MODELBEST_VOXCPM_MODEL_ID"), | ||
| base_url=os.getenv("MODELBEST_VOXCPM_BASE_URL", voxcpm.DEFAULT_BASE_URL), | ||
| ref_audio=reference_audio, | ||
| ) | ||
|
|
||
| started = time.perf_counter() | ||
| first_chunk_at: float | None = None | ||
| pcm_chunks: list[PcmData] = [] | ||
| try: | ||
| async for chunk in instance.send_iter(text): | ||
| if chunk.data is not None: | ||
| if first_chunk_at is None: | ||
| first_chunk_at = time.perf_counter() | ||
| pcm_chunks.append(chunk.data) | ||
| finally: | ||
| await instance.close() | ||
|
|
||
| if not pcm_chunks: | ||
| raise RuntimeError("VoxCPM returned no audio chunks") | ||
|
|
||
| pcm = b"".join(chunk.to_bytes() for chunk in pcm_chunks) | ||
| sample_rate = pcm_chunks[0].sample_rate | ||
| channels = pcm_chunks[0].channels | ||
| with wave.open(OUTPUT_PATH, "wb") as wav_file: | ||
| wav_file.setnchannels(channels) | ||
| wav_file.setsampwidth(2) | ||
| wav_file.setframerate(sample_rate) | ||
| wav_file.writeframes(pcm) | ||
|
|
||
| elapsed = time.perf_counter() - started | ||
| first_chunk_seconds = (first_chunk_at or started) - started | ||
| duration = len(pcm) / (sample_rate * channels * 2) | ||
| print(f"chunks: {len(pcm_chunks)}") | ||
| print(f"sample rate: {sample_rate} Hz") | ||
| print(f"time to first chunk: {first_chunk_seconds:.3f}s") | ||
| print(f"duration: {duration:.2f}s") | ||
| print(f"elapsed: {elapsed:.3f}s -> {OUTPUT_PATH}") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| asyncio.run(main()) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
|
|
||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,40 @@ | ||
| [build-system] | ||
| requires = ["hatchling", "hatch-vcs"] | ||
| build-backend = "hatchling.build" | ||
|
|
||
| [project] | ||
| name = "vision-agents-plugins-voxcpm" | ||
| dynamic = ["version"] | ||
| description = "ModelBest VoxCPM cloud TTS integration for Vision Agents" | ||
| readme = "README.md" | ||
| keywords = ["voxcpm", "modelbest", "TTS", "text-to-speech", "voice cloning", "voice agents"] | ||
| requires-python = ">=3.10" | ||
| license = "MIT" | ||
| dependencies = [ | ||
| "vision-agents", | ||
| "aiohttp>=3.13.3", | ||
| ] | ||
|
|
||
| [project.urls] | ||
| Documentation = "https://visionagents.ai/" | ||
| Website = "https://visionagents.ai/" | ||
| Source = "https://github.com/GetStream/Vision-Agents" | ||
|
|
||
| [tool.hatch.version] | ||
| source = "vcs" | ||
| raw-options = { root = "..", search_parent_directories = true, fallback_version = "0.0.0" } | ||
|
|
||
| [tool.hatch.build.targets.wheel] | ||
| packages = ["vision_agents"] | ||
|
|
||
| [tool.hatch.build.targets.sdist] | ||
| include = ["/vision_agents"] | ||
|
|
||
| [tool.uv.sources] | ||
| vision-agents = { workspace = true } | ||
|
|
||
| [dependency-groups] | ||
| dev = [ | ||
| "pytest>=8.4.1", | ||
| "pytest-asyncio>=1.0.0", | ||
| ] |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win
Move
py.typedinto the import package.The wheel only includes the
vision_agentstree. This root-level marker is not installed withvision_agents.plugins.voxcpm.Move it to
plugins/voxcpm/vision_agents/plugins/voxcpm/py.typed.