Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ agent = Agent(

**STT:** [Deepgram](https://visionagents.ai/integrations/deepgram) · [AssemblyAI](https://www.assemblyai.com/docs/streaming/universal-3-pro) · [Fast-Whisper](https://visionagents.ai/integrations/fast-whisper) · [Fish Audio](https://visionagents.ai/integrations/fish) · [Wizper](https://visionagents.ai/integrations/wizper) · [Mistral Voxtral](https://visionagents.ai/integrations/mistral) · [Telnyx](https://github.com/GetStream/Vision-Agents/tree/main/plugins/telnyx)

**TTS:** [ElevenLabs](https://visionagents.ai/integrations/elevenlabs) · [Cartesia](https://visionagents.ai/integrations/cartesia) · [Deepgram](https://visionagents.ai/integrations/deepgram) · [AWS Polly](https://visionagents.ai/integrations/aws-polly) · [Pocket](https://visionagents.ai/integrations/pocket) · [Kokoro](https://visionagents.ai/integrations/kokoro) · [Inworld](https://visionagents.ai/integrations/inworld) · [Fish Audio](https://visionagents.ai/integrations/fish) · [Telnyx](https://github.com/GetStream/Vision-Agents/tree/main/plugins/telnyx)
**TTS:** [ElevenLabs](https://visionagents.ai/integrations/elevenlabs) · [Cartesia](https://visionagents.ai/integrations/cartesia) · [Deepgram](https://visionagents.ai/integrations/deepgram) · [AWS Polly](https://visionagents.ai/integrations/aws-polly) · [Pocket](https://visionagents.ai/integrations/pocket) · [Kokoro](https://visionagents.ai/integrations/kokoro) · [Inworld](https://visionagents.ai/integrations/inworld) · [Fish Audio](https://visionagents.ai/integrations/fish) · [VoxCPM](https://github.com/GetStream/Vision-Agents/tree/main/plugins/voxcpm) · [Telnyx](https://github.com/GetStream/Vision-Agents/tree/main/plugins/telnyx)

**Vision:** [Ultralytics](https://visionagents.ai/integrations/ultralytics) · [Roboflow](https://visionagents.ai/integrations/roboflow) · [Moondream](https://visionagents.ai/integrations/moondream) · [TwelveLabs](https://github.com/GetStream/Vision-Agents/tree/main/plugins/twelvelabs) · [NVIDIA](https://visionagents.ai/integrations/nvidia) · [Decart](https://visionagents.ai/integrations/decart)

Expand Down
1 change: 1 addition & 0 deletions agents-core/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -89,6 +89,7 @@ tencent = ["vision-agents-plugins-tencent; sys_platform == 'linux'"]
minimax = ["vision-agents-plugins-minimax"]
twelvelabs = ["vision-agents-plugins-twelvelabs"]
speechify = ["vision-agents-plugins-speechify"]
voxcpm = ["vision-agents-plugins-voxcpm"]


[tool.hatch.metadata]
Expand Down
62 changes: 62 additions & 0 deletions plugins/voxcpm/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# VoxCPM Plugin

This package integrates the hosted [ModelBest VoxCPM](https://platform.modelbest.cn/console/docs/api/audio) Text-to-Speech API with Vision Agents. It streams audio as it is synthesized and supports optional voice cloning without requiring a local GPU.

## Installation

```bash
uv add "vision-agents[voxcpm]"
# or directly
uv add vision-agents-plugins-voxcpm
```

## Usage

Create a ModelBest API key and select a model with the `speech_synthesis` capability, then set:

```bash
export MODELBEST_API_KEY="your-api-key"
export MODELBEST_VOXCPM_MODEL_ID="your-model-id"
```

```python
from vision_agents.plugins import voxcpm

tts = voxcpm.TTS()

try:
async for chunk in tts.send_iter("Hello from VoxCPM!"):
if chunk.data:
print(chunk.data.duration_ms)
finally:
await tts.close()
```

## Voice cloning

`ref_audio` sets speaker identity. `prompt_audio` and its exact `prompt_text` can additionally preserve the reference delivery, including pacing, emotion, and pronunciation.

```python
tts = voxcpm.TTS(
ref_audio="speaker.wav",
prompt_audio="delivery.wav",
prompt_text="The exact words spoken in delivery.wav.",
)
```

Reference files must be valid, uncompressed WAV files no larger than 5 MiB. Audio supplied for cloning is sent to ModelBest; only use recordings you have permission to process and clone.

For a minimal live API check, see [`example/`](example/README.md). It writes the streamed response to a WAV file and can optionally test reference-audio cloning.

## Configuration

| Parameter | Default | Description |
|---|---|---|
| `api_key` | `MODELBEST_API_KEY` | ModelBest API key. |
| `model` | `MODELBEST_VOXCPM_MODEL_ID` | Model ID with `speech_synthesis` capability. |
| `voice` | `"default"` | Protocol voice value; cloned identity comes from `ref_audio`. |
| `base_url` | `https://api.modelbest.cn/v1` | ModelBest API base URL. |
| `ref_audio` | `None` | WAV bytes or path used for speaker identity. |
| `prompt_audio` | `None` | WAV bytes or path used for delivery cloning. |
| `prompt_text` | `None` | Exact transcript paired with `prompt_audio`. |
| `request_timeout` | `120.0` | Maximum seconds without response data. |
5 changes: 5 additions & 0 deletions plugins/voxcpm/example/.env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
MODELBEST_API_KEY=your-modelbest-api-key
MODELBEST_VOXCPM_MODEL_ID=your-speech-synthesis-model-id
MODELBEST_VOXCPM_BASE_URL=https://api.modelbest.cn/v1
# Optional: path to a PCM16 WAV for speaker cloning.
VOXCPM_REFERENCE_AUDIO=
21 changes: 21 additions & 0 deletions plugins/voxcpm/example/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# ModelBest VoxCPM Smoke Test

This example makes one real request to the hosted ModelBest VoxCPM API and writes the returned speech to a WAV file. It is the smallest way to verify API credentials, model access, streamed audio parsing, and optional speaker cloning before adding VoxCPM to a full Agent.

## Setup

```bash
cd plugins/voxcpm/example
cp .env.example .env
```

Fill in `MODELBEST_API_KEY` and a ModelBest model ID with the `speech_synthesis` capability. Set `VOXCPM_REFERENCE_AUDIO` to a PCM16 WAV path to test speaker cloning.

## Run

```bash
uv run voxcpm_smoke.py
uv run voxcpm_smoke.py "你好,这是 VoxCPM 的测试语音。"
```

The output is written to `voxcpm_smoke.wav`. The command prints the number of streamed chunks, sample rate, first-chunk latency, and total duration.
Empty file.
13 changes: 13 additions & 0 deletions plugins/voxcpm/example/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
[project]
name = "voxcpm-tts-example"
version = "0.0.0"
requires-python = ">=3.10"

dependencies = [
"python-dotenv>=1.0",
"vision-agents-plugins-voxcpm",
]

[tool.uv.sources]
"vision-agents-plugins-voxcpm" = { path = "..", editable = true }
"vision-agents" = { path = "../../../agents-core", editable = true }
67 changes: 67 additions & 0 deletions plugins/voxcpm/example/voxcpm_smoke.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
"""Make one hosted VoxCPM request and write the streamed audio to WAV."""

import asyncio
import os
import sys
import time
import wave

from dotenv import load_dotenv
from getstream.video.rtc import PcmData
from vision_agents.plugins import voxcpm

load_dotenv()

OUTPUT_PATH: str = "voxcpm_smoke.wav"


async def main() -> None:
text = (
sys.argv[1]
if len(sys.argv) > 1
else "Hello from the Vision Agents VoxCPM plugin."
)
reference_audio = os.getenv("VOXCPM_REFERENCE_AUDIO") or None
instance = voxcpm.TTS(
api_key=os.getenv("MODELBEST_API_KEY"),
model=os.getenv("MODELBEST_VOXCPM_MODEL_ID"),
base_url=os.getenv("MODELBEST_VOXCPM_BASE_URL", voxcpm.DEFAULT_BASE_URL),
ref_audio=reference_audio,
)

started = time.perf_counter()
first_chunk_at: float | None = None
pcm_chunks: list[PcmData] = []
try:
async for chunk in instance.send_iter(text):
if chunk.data is not None:
if first_chunk_at is None:
first_chunk_at = time.perf_counter()
pcm_chunks.append(chunk.data)
finally:
await instance.close()

if not pcm_chunks:
raise RuntimeError("VoxCPM returned no audio chunks")

pcm = b"".join(chunk.to_bytes() for chunk in pcm_chunks)
sample_rate = pcm_chunks[0].sample_rate
channels = pcm_chunks[0].channels
with wave.open(OUTPUT_PATH, "wb") as wav_file:
wav_file.setnchannels(channels)
wav_file.setsampwidth(2)
wav_file.setframerate(sample_rate)
wav_file.writeframes(pcm)

elapsed = time.perf_counter() - started
first_chunk_seconds = (first_chunk_at or started) - started
duration = len(pcm) / (sample_rate * channels * 2)
print(f"chunks: {len(pcm_chunks)}")
print(f"sample rate: {sample_rate} Hz")
print(f"time to first chunk: {first_chunk_seconds:.3f}s")
print(f"duration: {duration:.2f}s")
print(f"elapsed: {elapsed:.3f}s -> {OUTPUT_PATH}")


if __name__ == "__main__":
asyncio.run(main())
1 change: 1 addition & 0 deletions plugins/voxcpm/py.typed
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Move py.typed into the import package.

The wheel only includes the vision_agents tree. This root-level marker is not installed with vision_agents.plugins.voxcpm.

Move it to plugins/voxcpm/vision_agents/plugins/voxcpm/py.typed.

40 changes: 40 additions & 0 deletions plugins/voxcpm/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
[build-system]
requires = ["hatchling", "hatch-vcs"]
build-backend = "hatchling.build"

[project]
name = "vision-agents-plugins-voxcpm"
dynamic = ["version"]
description = "ModelBest VoxCPM cloud TTS integration for Vision Agents"
readme = "README.md"
keywords = ["voxcpm", "modelbest", "TTS", "text-to-speech", "voice cloning", "voice agents"]
requires-python = ">=3.10"
license = "MIT"
dependencies = [
"vision-agents",
"aiohttp>=3.13.3",
]

[project.urls]
Documentation = "https://visionagents.ai/"
Website = "https://visionagents.ai/"
Source = "https://github.com/GetStream/Vision-Agents"

[tool.hatch.version]
source = "vcs"
raw-options = { root = "..", search_parent_directories = true, fallback_version = "0.0.0" }

[tool.hatch.build.targets.wheel]
packages = ["vision_agents"]

[tool.hatch.build.targets.sdist]
include = ["/vision_agents"]

[tool.uv.sources]
vision-agents = { workspace = true }

[dependency-groups]
dev = [
"pytest>=8.4.1",
"pytest-asyncio>=1.0.0",
]
Loading
Loading