Repository navigation
Expand file tree
/
Copy pathdocs.json
More file actions
188 lines (188 loc) · 133 KB
/
Copy pathdocs.json
File metadata and controls
188 lines (188 loc) · 133 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
{
"project": "vox",
"version": "0.1.0",
"generatedBy": "scripts/generate-agent-artifacts.ts",
"sections": [
{
"id": "start",
"title": "Start here",
"description": "Pick the shortest path to working speech with Vox, in your app, on your Mac, from Node, or from the browser.",
"content": "Vox is open-source speech-to-text and text-to-speech for Mac apps. Transcription runs on your Mac. Speech output uses the system voice on device, or a cloud voice when you bring a key.\n\nPick the row that matches what you are building. Each path ends with something you can run.\n\n| You are building | Start with | Runs where | First result |\n|---|---|---|---|\n| Just want to try it | [Try Minivox](./start-mac.md#try-minivox-in-60-seconds) | macOS 26 | Dictate anywhere, paste the text |\n| A Swift app | [Vox in your app](./start-swift.md) | Inside your app | `VoxDictation` start, stop, text |\n| A Bun or Node tool | [Vox for Node](./start-node.md) | Talks to Vox on your Mac | Transcribe a file, speak a line |\n| A web page or extension | [Vox for the browser](./start-browser.md) | Talks to Vox on your Mac | Dictate into a page |\n| Anything, with a coding agent | [Build with an agent](./start-agent.md) | Wherever the agent works | A prompt that wires Vox in |\n\n## Two ways Vox runs\n\n- **Vox in your app.** Your Swift app links the Vox packages and runs speech in its own process. Nothing else to install for your users.\n- **Vox on your Mac.** The Vox app runs a local engine that other programs talk to: Node and Bun tools over WebSocket, web pages over a local HTTP bridge. Your users install Vox once; every tool on the Mac shares one loaded model.\n\nBoth keep the same habits:\n\n- **Preload.** Loading a model takes seconds the first time. Vox makes that step explicit (`warmUp`, `preloadModel`) so you can do it when the user shows intent, not on their first word.\n- **Timings.** Every transcription can record how long each stage took, tagged with your app's name, so you can see where time goes. See [Observability](./observability.md).\n\n## Requirements\n\n- An Apple Silicon Mac. Intel Macs are not supported.\n- macOS 14 or later for Vox in your app; macOS 26 or later for Minivox and the Vox app.\n- iOS 17 or later for file transcription inside an iOS app.\n- Bun 1.2 or later, or Node 22 or later, for the TypeScript packages.\n\n## What Vox does not do yet\n\nSaying this up front saves you an afternoon:\n\n- **No microphone capture on iOS.** On iOS, record audio in your app and hand Vox the file. Microphone capture in `VoxDictation` is macOS only.\n- **No live partial text inside your app.** In-app dictation returns the transcript when you stop. Live partials are available through Vox on your Mac (the browser and Node live sessions).\n- **Speech output is not all local.** The default high-quality voice, `gpt-4o-mini-tts`, calls OpenAI and needs your API key. The local option is the system voice, `avspeech:system`. Vox does not switch between them for you.\n- **The browser client only listens.** `@voxd/client` transcribes and aligns. To speak text, use Node, the CLI, or Swift.\n\n## Reference\n\nOnce something runs, these pages hold the details:\n\n- [Swift embed reference](./apple-embed.md): every embed type, speech output, packaging.\n- [Node SDK reference](./sdk.md) and [Browser client reference](./web-integration.md).\n- [Command line](./quickstart.md): install, health checks, benchmarks.\n- [Models and plugins](./models.md), [Providers](./providers.md), [Observability](./observability.md)."
},
{
"id": "start-mac",
"title": "Vox on your Mac",
"description": "Try Minivox in sixty seconds, then install the Vox app so Node tools and web pages can use speech on your Mac.",
"content": "## Try Minivox in 60 seconds\n\nMinivox is the smallest dictation app built on Vox. It needs macOS 26.\n\n```bash\nbunx @voxd/cli@latest install mini\n```\n\n1. Put the text cursor wherever you want the words.\n2. Press **Right ⌘M** and allow microphone and Accessibility access.\n3. Speak, then press **Right ⌘M** again. The text is copied and pasted where your cursor was.\n\nThe first dictation downloads the speech model, about 500 MB, so expect a pause once. Run `minivox settings` to change the shortcut or microphone.\n\nMinivox uses [Vox in your app](./start-swift.md): there is no daemon, just a Swift app linking the Vox packages. Its source is in [`apps/minivox`](https://github.com/hudsonkit/vox/tree/main/apps/minivox) and short enough to read in one sitting.\n\n## Install the Vox app\n\nInstall Vox on your Mac when other programs should share one speech engine: Node and Bun tools, web pages, browser extensions, the command line.\n\n1. Download Vox from [voxd.cc/download](https://voxd.cc/download) and drag it to Applications.\n2. Open it. Vox lives in the menu bar and starts the local engine.\n3. Check it from a terminal:\n\n```bash\nbunx @voxd/cli@latest doctor # expect ready: true\n```\n\nThe app runs two local endpoints, both on `127.0.0.1` only:\n\n| Endpoint | Used by | Default port |\n|---|---|---|\n| WebSocket engine | `@voxd/sdk`, the CLI | `42137` |\n| HTTP bridge | `@voxd/client` in web pages | `43115` |\n\nThe HTTP bridge runs inside the Vox app. If only the engine is running, for example after `vox daemon start`, Node tools work but web pages will not find Vox.\n\n## Try it from the terminal\n\n```bash\nalias vox=\"bunx @voxd/cli@latest\"\nvox warmup start parakeet:v3 # preload the model\nvox transcribe file --metrics recording.wav # text plus stage timings\nvox speak --model avspeech:system \"Hello\" # system voice, no key needed\nvox perf dashboard # timings by app and model\n```\n\n## Next\n\n- Use it from code: [Vox for Node](./start-node.md) or [Vox for the browser](./start-browser.md).\n- All commands, benchmarks and troubleshooting: [Command line](./quickstart.md)."
},
{
"id": "start-swift",
"title": "Vox in your app",
"description": "Add on-device dictation and speech output to a Swift app with VoxDictation, from package setup to a shippable app.",
"content": "Vox in your app means your Swift app links the Vox packages and runs speech in its own process. Your users install nothing else.\n\nYou will end up with dictation in four calls: `warmUp()`, `start()`, `stop()` for text, or `cancel()`.\n\n## 1. Add the package\n\nIn Xcode, choose **File → Add Package Dependencies…**, enter `https://github.com/hudsonkit/vox`, and add the `VoxCore` and `VoxEngine` libraries to your app target.\n\nOr in `Package.swift`:\n\n```swift\ndependencies: [\n .package(url: \"https://github.com/hudsonkit/vox.git\", from: \"0.5.2\"),\n],\ntargets: [\n .target(name: \"MyApp\", dependencies: [\n .product(name: \"VoxCore\", package: \"vox\"),\n .product(name: \"VoxEngine\", package: \"vox\"),\n ]),\n]\n```\n\nVox needs macOS 14 or iOS 17, on Apple Silicon.\n\n## 2. Dictate\n\n```swift\nimport VoxCore\nimport VoxEngine\n\nlet dictation = VoxDictation(clientId: \"my-app\")\n\n// When the user shows intent, for example opening the compose view:\ntry await dictation.warmUp()\n\n// Mic button down:\ntry await dictation.start()\n\n// Mic button up:\nlet result = try await dictation.stop()\nprint(result.text)\n```\n\n- `warmUp()` downloads the speech model the first time (about 500 MB) and loads it into memory. Later launches only load it. Call it before the user needs it; if you don't, the first `stop()` waits for it. Pass a closure to show progress: `warmUp { progress in … }`.\n- `start()` records the default microphone to a temporary file. `stop()` transcribes it, deletes the file and returns a `TranscriptionOutput` with `text`, `words` timings and `metrics`.\n- `cancel()` stops and throws the audio away. `inputLevel()` gives a 0 to 1 level for a meter while recording.\n- `start(onBuffer:)` hands you live 16 kHz audio buffers if you want your own waveform.\n\nAlready have audio? `try await dictation.transcribe(fileURL: url)` works on macOS and iOS.\n\nThe default model is `parakeet:v3`, which handles 25 European languages. `VoxDictation(clientId:modelId: \"parakeet:v2\")` uses the English-only model.\n\n## 3. Ask for the microphone\n\nAdd `NSMicrophoneUsageDescription` to your Info.plist with a sentence your users will see, for example *\"Dictation turns your voice into text on this Mac.\"*\n\nA sandboxed Mac app also needs the **Audio Input** capability (`com.apple.security.device.audio-input`) and **Outgoing Connections** (`com.apple.security.network.client`) for the first model download.\n\n## 4. Speak text\n\nSpeech output is a separate, optional piece. Add the `VoxAppleSpeech` library, then keep one controller per place in your app that talks:\n\n```swift\nimport VoxCore\nimport VoxEngine\nimport VoxAppleSpeech\n\nlet speech = AppleSpeechOutputController(onEvent: { event in\n print(event.phase) // .generating, .playing, .finished, .cancelled, .failed\n})\n\nawait speech.speak(SynthesisRequest(text: \"Your draft is saved.\", modelId: TTSDefaults.localModelId))\nawait speech.stop()\n```\n\nA new `speak` replaces whatever was playing. `TTSDefaults.localModelId` is `avspeech:system`, the built-in system voice: on device, free, no key.\n\nFor a more natural voice, use `gpt-4o-mini-tts` (`TTSDefaults.modelId`). It calls OpenAI with your API key, so give the engine the key explicitly:\n\n```swift\nlet engine = TTSEngineManager(provider: TTSProviderRegistry(config: ProvidersConfig(providers: [\n ProviderEntry(id: \"avspeech\", kind: .tts, builtin: true,\n models: [AVSpeechSynthesizerProvider.modelID]),\n ProviderEntry(id: \"openai-tts\", kind: .tts, builtin: true,\n models: OpenAITTSProvider.supportedModelIDs,\n env: [\"OPENAI_API_KEY\": apiKey]),\n])))\nlet speech = AppleSpeechOutputController(engine: engine)\nawait speech.speak(SynthesisRequest(text: \"Your draft is saved.\", modelId: TTSDefaults.modelId))\n```\n\nVox does not fall back from OpenAI to the system voice on its own. If there is no key or no network, the request fails with `.failed`; choose the model in your app. Don't ship an OpenAI key inside an app binary: fetch a short-lived key from your server, or let users bring their own.\n\n## 5. See where time goes\n\n`VoxDictation` records the timings of every transcription to `~/.vox/performance.jsonl` on macOS, or `Application Support/Vox/performance.jsonl` on iOS, tagged with your `clientId`, the route (`transcribe.dictation` or `transcribe.file`) and the model. With the Vox command line installed:\n\n```bash\nvox perf dashboard --client my-app\n```\n\nPass `recordsPerformance: false` to turn this off.\n\n## 6. Ship it\n\n- **Resource bundle.** Vox's model catalog lives in a SwiftPM resource bundle. Xcode copies it for you. If you package the app yourself, copy the bundle SwiftPM built, `Vox_HudsonSpeechEngine.bundle` or `HudsonSpeechEngine_HudsonSpeechEngine.bundle` depending on which package you depend on, into `YourApp.app/Contents/Resources`. A packaged app never falls back to a developer build directory.\n- **Model download.** The first `warmUp()` downloads from Hugging Face. Ship with outgoing network allowed, and show progress.\n- **Microphone and sandbox.** The usage string and entitlements from step 3.\n- **Signing.** Nothing Vox-specific; sign and notarize as usual.\n\n## On iOS\n\nOn iOS, `VoxDictation` preloads and transcribes files; `start()` is macOS only for now. Record with `AVAudioRecorder` or `AVAudioEngine` to a file, then call `transcribe(fileURL:)`. On iOS the app must also set up its own `AVAudioSession` for recording.\n\n## A runnable example\n\n[`examples/swift-embed`](https://github.com/hudsonkit/vox/tree/main/examples/swift-embed) is a small command-line program built on `VoxDictation`:\n\n```bash\ncd examples/swift-embed\nswift run vox-embed-demo listen 5\nswift run vox-embed-demo speak \"Hello from Vox\"\n```\n\n[Minivox](https://github.com/hudsonkit/vox/tree/main/apps/minivox) is a real menu-bar app built from the same pieces.\n\n## Next\n\n- Every type behind `VoxDictation`, other speech providers, and the speech output controller in depth: [Swift embed reference](./apple-embed.md).\n- Other models: [Models and plugins](./models.md)."
},
{
"id": "start-node",
"title": "Vox for Node",
"description": "Transcribe, dictate and speak from a Bun or Node tool through Vox on your Mac.",
"content": "`@voxd/sdk` lets a Bun or Node program use the speech engine in the Vox app on the same Mac. The model stays loaded across every tool that uses it.\n\n## 1. Run Vox on your Mac\n\nInstall and open the Vox app, then check it: see [Vox on your Mac](./start-mac.md#install-the-vox-app).\n\n```bash\nbunx @voxd/cli@latest doctor # expect ready: true\n```\n\n## 2. Install the SDK\n\n```bash\nbun add @voxd/sdk # or: npm install @voxd/sdk\n```\n\n## 3. Transcribe a file and speak a line\n\n```ts\nimport { writeFile } from \"node:fs/promises\";\nimport { VoxClient } from \"@voxd/sdk\";\n\nconst vox = new VoxClient({ clientId: \"my-tool\" });\nawait vox.connect();\n\nawait vox.preloadModel(\"parakeet:v3\");\nconst result = await vox.transcribeFile(\"/absolute/path/to/audio.wav\", \"parakeet:v3\");\nconsole.log(result.text, `${Math.round(result.elapsedMs)} ms`);\n\nconst speech = await vox.synthesize(\"Hello from Vox.\", { modelId: \"avspeech:system\", format: \"wav\" });\nawait writeFile(\"hello.wav\", speech.audio);\n\nvox.disconnect();\n```\n\n- `clientId` names your tool in Vox's timings. Pick something stable.\n- `preloadModel` loads the model before the first request. Skip it and the first transcription waits for the load.\n- File paths must be absolute: Vox reads the file itself.\n- `avspeech:system` is the Mac's system voice. For `gpt-4o-mini-tts`, add an OpenAI key in the Vox app, or pass `credentials: { OPENAI_API_KEY }` in the options.\n\n## 4. Dictate with live text\n\nVox records from the Mac's microphone and streams partial text while the user speaks:\n\n```ts\nconst session = vox.createLiveSession();\nsession.on(\"partial\", ({ text }) => process.stdout.write(`\\r${text}`));\n\nconst done = session.start();\nsetTimeout(() => session.stop(), 5000);\nconsole.log(\"\\n\" + (await done).text);\n```\n\nThe first live session makes macOS ask whether Vox may use the microphone.\n\n## Runnable example\n\n[`examples/node-hello`](https://github.com/hudsonkit/vox/tree/main/examples/node-hello):\n\n```bash\ncd examples/node-hello\nbun install\nbun run index.ts path/to/audio.wav\n```\n\n## Next\n\n- Every method, the result shapes and error codes: [Node SDK reference](./sdk.md).\n- Timings by tool and model: `vox perf dashboard --client my-tool`, see [Observability](./observability.md)."
},
{
"id": "start-browser",
"title": "Vox for the browser",
"description": "Add dictation and file transcription to a web page or extension through Vox on your Mac.",
"content": "`@voxd/client` lets a web page or browser extension use Vox on the visitor's Mac through a small HTTP bridge on `127.0.0.1`. No server of yours handles the audio.\n\nThe browser client listens and transcribes. To speak text, use [Node](./start-node.md), the command line, or [Swift](./start-swift.md).\n\n## 1. Run Vox on your Mac\n\nInstall and open the Vox app: see [Vox on your Mac](./start-mac.md#install-the-vox-app). The bridge runs inside the app, so the app must be open, not only the background engine.\n\n## 2. Allow your page's origin\n\nVox only answers pages it trusts. Any `http://localhost` or `http://127.0.0.1` port is allowed out of the box, so local development needs no setup.\n\nFor your own domain, add it in the Vox app's settings, or drop a JSON file into `~/.vox/origins.d/`:\n\n```json\n{\"origins\":[\"https://app.example.com\"]}\n```\n\n## 3. Install the client\n\n```bash\nbun add @voxd/client # or: npm install @voxd/client\n```\n\n## 4. Find Vox, then dictate\n\n```ts\nimport { createVoxdClient } from \"@voxd/client\";\n\nconst vox = createVoxdClient({ clientId: \"my-site\" });\n\nif (await vox.probe()) {\n const session = vox.createLiveSession();\n session.onPartial(({ text }) => { output.textContent = text; });\n\n stopButton.onclick = () => session.stop();\n const final = await session.start();\n output.textContent = final.text;\n}\n```\n\n- `probe()` returns `false` quickly when Vox is not running, so it is safe on every page load. Offer a link to [voxd.cc/download](https://voxd.cc/download) in that case.\n- Vox records from the Mac's microphone, so your page needs no microphone permission of its own.\n- An origin that is not allowed fails on the first call after `probe()`, such as `capabilities()` or `start()`.\n\n## Transcribe a file or recording\n\n```ts\nconst result = await vox.transcribe({ audio: file, timestamps: true });\nresult.text; // full transcript\nresult.words; // [{ word, start, end }, ...]\n```\n\n`audio` takes a `Blob`, `File` or `ArrayBuffer`.\n\n## Runnable example\n\n[`examples/web-hello`](https://github.com/hudsonkit/vox/tree/main/examples/web-hello) is one HTML page with dictation and file upload:\n\n```bash\ncd examples/web-hello\nbun install\nbun run start # http://localhost:3000\n```\n\n## Next\n\n- Alignment jobs, fallbacks, error codes and the bridge endpoints: [Browser client reference](./web-integration.md)."
},
{
"id": "start-agent",
"title": "Build with an agent",
"description": "Copy-paste prompts that get a coding agent to add Vox to your project correctly the first time.",
"content": "Coding agents do well with Vox when they read the right page first. Each prompt below points the agent at one page and states the facts it most often gets wrong.\n\nEvery page on this site has a compact agent version, and the whole set is at [voxd.cc/llms.txt](https://voxd.cc/llms.txt) and [voxd.cc/llms-full.txt](https://voxd.cc/llms-full.txt).\n\n## Add dictation to a Swift app\n\n```text\nAdd on-device dictation to this app with Vox. Read https://voxd.cc/docs/start-swift first and follow it.\n\n- Add the Swift package https://github.com/hudsonkit/vox (products VoxCore and VoxEngine).\n- Use VoxDictation. Call warmUp() when the user opens the screen that dictates, not at app launch.\n- Wire a mic button: start() on press, stop() on release, put result.text into the text field.\n- Add NSMicrophoneUsageDescription to Info.plist.\n- On iOS, start() is not available: record to a file and call transcribe(fileURL:).\n- Keep clientId stable; it names this app in Vox's timings.\n```\n\n## Add spoken replies to a Swift app\n\n```text\nAdd spoken output to this app with Vox. Read https://voxd.cc/docs/start-swift#4-speak-text first.\n\n- Add the VoxAppleSpeech product from https://github.com/hudsonkit/vox.\n- Use one AppleSpeechOutputController per place in the UI that speaks.\n- Default to the system voice, TTSDefaults.localModelId. Only use gpt-4o-mini-tts if the user supplies an OpenAI key; Vox does not fall back between them automatically.\n- Never hard-code an API key in the app.\n```\n\n## Use Vox from a Bun or Node tool\n\n```text\nUse Vox on this Mac for speech in this tool. Read https://voxd.cc/docs/start-node first.\n\n- Install @voxd/sdk and connect with new VoxClient({ clientId: \"<this tool's name>\" }).\n- Call preloadModel(\"parakeet:v3\") before the first transcription.\n- Pass absolute file paths to transcribeFile.\n- If connect() fails, tell the user to open the Vox app (https://voxd.cc/download).\n```\n\n## Add dictation to a web page\n\n```text\nAdd dictation to this web page using Vox on the visitor's Mac. Read https://voxd.cc/docs/start-browser first.\n\n- Install @voxd/client and create the client with a stable clientId.\n- Call probe() on load; if false, show a link to https://voxd.cc/download and hide the mic button.\n- Use createLiveSession(): onPartial for live text, start() to begin, stop() to finish.\n- localhost works out of the box. For production, the origin must be added in Vox settings or ~/.vox/origins.d/.\n- The browser client does not speak text.\n```\n\n## What agents get wrong without these\n\n- Inventing a `VoxClient` in Swift. Swift apps use `VoxDictation` and the Swift packages, never the TypeScript SDK.\n- Calling `warmUp()` at launch, which costs memory before the user asks for anything, or never calling it, which makes the first dictation slow.\n- Assuming the OpenAI voice falls back to the system voice.\n- Assuming iOS can record through Vox.\n- Using the old `github.com/arach/vox` URL. The repository is `github.com/hudsonkit/vox`."
},
{
"id": "overview",
"title": "Vox Overview",
"description": "Choose the right Vox integration mode and understand the stack's runtime, package, and telemetry surfaces.",
"content": "Vox is a local-first voice stack for macOS and iOS. It supports both speech-to-text (STT / ASR) and text-to-speech (TTS), and you can use it in two main ways:\n\n- Embed mode: Apple apps link Vox's Swift packages directly and keep speech in and speech out in process.\n- Companion mode: `voxd` exposes the runtime over local WebSocket JSON-RPC, and `VoxBridge` / `voxbridge` exposes an HTTP bridge for browser clients when web or shared-process access is useful.\n\nMain surfaces:\n\n- Swift packages: `VoxCore` and `VoxEngine` (start with `VoxDictation`), plus optional `VoxAppleSpeech` for speech output, for apps that run Vox in process. `VoxService` and `VoxBridge` are the companion runtime itself, not needed in an app.\n- `voxd`: Vox Companion, the Swift daemon. Warm-up, telemetry, bridge transport, shared-process coordination.\n- `@voxd/sdk`: TypeScript SDK for Bun/Node and other companion-connected integrations. WebSocket JSON-RPC to `voxd`.\n- `@voxd/client`: Browser SDK. HTTP bridge to the Vox Companion for web apps.\n- `@voxd/cli`: Node CLI. Health checks, model management, transcription, synthesis, voices, warm-up, and benchmarks.\n\n## Which surface to reach for\n\n- Use the Swift packages when you are modifying a macOS or iOS app and want speech to stay in process.\n- Use `@voxd/sdk` when you want a local Bun or Node client to talk to the companion daemon over WebSocket JSON-RPC.\n- Use `@voxd/client` when you want a browser app to talk to the local HTTP bridge on the same Mac.\n- Use `@voxd/cli` when you want operator commands, warm-up controls, voice listing, or reproducible benchmarks.\n\n## Why it exists\n\nMany voice stacks hide lifecycle, warm-up, and latency. Vox tries to keep those parts visible:\n\n- Model stays local\n- STT and TTS stay explicit runtime capabilities instead of hidden backend switches\n- Warm-up is an explicit API\n- Latency dimensions (`clientId`, `route`, `modelId`) are preserved\n- Runtime stays easy to inspect from the start\n\n## Repository layout\n\n- `swift/`: VoxCore, VoxEngine, VoxService, VoxBridge, voxd companion\n- `apps/vox/`: full Vox companion app\n- `apps/minivox/`: small, single-purpose menu-bar dictation app\n- `packages/client/`: `@voxd/sdk` (TypeScript SDK)\n- `packages/web-client/`: `@voxd/client` (browser SDK)\n- `packages/cli/`: `@voxd/cli` (Node CLI)\n- `docs/`: Dewey source content\n- `site/`: website and docs UI\n\n## Design principles\n\n1. Root cause over workaround.\n2. Warm-up is part of the product, not an implementation detail.\n3. Instrumentation is part of the API surface.\n4. Multi-client support stays visible in the protocol and telemetry.\n\n## How it fits together\n\nApple app teams embed the Swift packages directly and keep the same provider, warm-up, and telemetry semantics in process. Bun and Node tools use `@voxd/sdk` against `voxd`, browser apps use `@voxd/client` against the local HTTP bridge, and operators use `vox` to check health, list voices, transcribe, synthesize, warm models, and benchmark both speech paths. Companion mode lets `voxd` stay warm across browser integrations, local tools, and other shared-process clients while the bridge stays narrow and browser-facing.\n\n## Reference implementations\n\n- `apps/minivox/` is the small, single-purpose menu-bar dictation app built on the direct Apple embed path.\n- `packages/client/` and `packages/web-client/` are good companion-mode SDK references in the repo.\n\nInstall the signed and notarized Minivox app through the existing CLI or Homebrew:\n\n```bash\nnpx -y @voxd/cli@latest install mini\n# or\nbrew install --cask arach/vox/minivox\n```\n\nThe installer opens Minivox automatically. Look for the waveform in the menu bar, put the text cursor where you want your dictation, and press **Right ⌘M** to start. Allow microphone access if asked, then press **Right ⌘M** again to stop. Minivox copies the result and pastes it when Accessibility access is enabled. The first dictation may download Parakeet.\n\nBoth installers expose a `minivox` command. Run `minivox settings` to change the shortcut or microphone, or `minivox quit` to stop the menu-bar app. The npm installer accepts `--quiet` and `--verbose` to control setup output.\n\n## Workflows\n\nThese examples assume `vox` is on your `PATH`. In a repo checkout, replace `vox` with `node packages/cli/dist/index.js` after `bun run build`.\n\n```bash\nbun install && bun run build\nvox daemon start\nvox doctor\n\nvox warmup start parakeet:v3\nvox transcribe file --model parakeet:v3 --metrics --timestamps /tmp/sample.wav\nvox models catalog\n\nvox voices --model gpt-4o-mini-tts\nvox speak --model gpt-4o-mini-tts --metrics \"Hello from Vox\"\n\nvox transcribe bench --model parakeet:v3 /tmp/sample.wav 5\nvox speak bench --model gpt-4o-mini-tts \"Hello from Vox\" 5\nvox perf dashboard --client vox-cli\n```\n\nContinue with the [Quickstart](./quickstart.md), [Models and plugins](./models.md), [Swift Embed Guide](./apple-embed.md), or [Web Integration Guide](./web-integration.md) for the surface you chose."
},
{
"id": "quickstart",
"title": "Command line",
"description": "Install Vox on your Mac, check its health, and exercise transcription, speech, preload and timings from the terminal.",
"content": "## Prerequisites\n\n- macOS 14+ or iOS 17+ for direct Swift transcription embedding; macOS 26+ for Minivox and the Vox menu app\n- Node 22+\n- Vox Companion installed from the DMG, or `voxd` available at `~/.vox/bin/voxd`\n- Swift 6.2+ only if you plan to build Vox from a repo checkout\n\n## Install and verify\n\n```bash\nbun add -g @voxd/cli # or: npm install -g @voxd/cli\nvox install\nvox doctor # expect ready: true\n```\n\nTo run a command without installing, use `bunx @voxd/cli@latest <command>`.\n\n`vox install` registers the LaunchAgent for an existing `voxd` binary. The simplest path is to install `Vox.dmg` first, then run the CLI.\n\nIf you are running from a repo checkout instead of a global install, replace `vox` with `node packages/cli/dist/index.js` after `bun run build`.\n\nIf you are building a browser client, pair the local companion with `@voxd/client` and the [Web Integration Guide](./web-integration.md).\n\n## Speech to text\n\n```bash\nvox warmup start parakeet:v3\nvox transcribe file --model parakeet:v3 /path/to/audio.wav --metrics --timestamps\nvox transcribe bench --model parakeet:v3 /path/to/audio.wav 5\nvox models catalog\n```\n\nWarm-up skips cold-start cost. `transcribe file` prints transcript text, stage timings, and optional word-level timestamps. `bench` gives you warm-path variance for the same clip. `parakeet:v3` stays the default; `parakeet:v2` is the English-only TDT option. `vox models catalog` shows the published dictation catalog. Gemma 4 and other new families install as plugins: `vox plugins install mlx-vlm`.\n\n## Text to speech\n\n```bash\nvox voices --model gpt-4o-mini-tts\nvox speak --model gpt-4o-mini-tts --metrics \"Hello from Vox\"\nvox speak bench --model gpt-4o-mini-tts \"Hello from Vox\" 5\n```\n\n`voices` shows available presets for the selected model. `speak` synthesizes audio immediately and prints synthesis metrics when `--metrics` is set. `speak bench` repeats the same request so you can compare warm-path TTS behavior.\n\n## External providers\n\nFor non-Parakeet ASR or non-default TTS, add entries to `~/.vox/providers.json` and then pass the returned model ID with `--model`.\n\nThe [Provider Protocol](./providers.md) includes built-in OpenAI, ElevenLabs, MiniMax, and `mlx-audio` examples for TTS.\n\n## Measure and inspect\n\n```bash\nvox perf dashboard --client vox-cli\nvox logs daemon --tail 80\nvox transcribe status\n```\n\n`perf dashboard` shows latency samples by client, route, and model. Use `logs daemon` and `transcribe status` when a live session gets stuck or the mic is busy.\n\n## Common failure cases\n\n- Missing ASR model: `vox models list` then `vox models install`\n- TTS provider or voice issue: `vox voices --model <id>` then retry `vox speak --model <id> ...`\n- External TTS model install/setup: follow the provider's own setup flow, such as the `mlx-audio` environment in `~/.vox/providers.json`\n- Wrong or missing voice: `vox voices --model <id>`\n- External provider missing dependencies: verify `~/.vox/providers.json` and any referenced interpreter or API key\n- Cold runtime: `vox warmup start` or `vox warmup schedule`\n- No performance data: run a `transcribe` or `speak` command first so the runtime emits samples\n- Stuck live session: `vox transcribe status` then `vox transcribe cancel`\n- Need daemon logs: `vox logs daemon --tail 120`\n\n## Next steps\n\nIf you are integrating Vox into a macOS or iOS app, start with [Vox in your app](./start-swift.md).\n\nIf you are choosing a dictation model or installing Gemma, read [Models and plugins](./models.md).\n\nIf you are wiring external STT or TTS engines into Vox Companion, read the [Provider Protocol](./providers.md).\n\nTry [Minivox](https://github.com/hudsonkit/vox/tree/main/apps/minivox), the small menu-bar dictation app built directly on Vox. The [transcribe TUI](https://github.com/hudsonkit/vox/tree/main/examples/transcribe-tui) remains the companion-connected terminal sample."
},
{
"id": "apple-embed",
"title": "Swift embed reference",
"description": "Every type for running Vox inside a macOS or iOS app, covering dictation, transcription, speech output, timings and packaging.",
"content": "This is the reference for running Vox inside a macOS or iOS app. New here? [Vox in your app](./start-swift.md) gets dictation working in a few minutes; come back for the details.\n\n## Choose the integration mode\n\n- Use embed mode when the caller is app code running inside a macOS or iOS process.\n- Use Vox Companion (`voxd`) when the caller lives outside the app process, such as a web app, browser extension, or Bun/Node tool.\n- Keep `voxd` out of the Apple app itself. If the goal is in-process app integration, embed the Swift packages directly instead.\n\n## What embed mode is today\n\nStart with `VoxDictation`. It wraps the pieces below in four calls: `warmUp()`, `start()`, `stop()`, `cancel()`, plus `transcribe(fileURL:)`. It records timings for you. Reach for the lower-level types when you need something it does not do.\n\n- dictation: `VoxDictation` (microphone capture on macOS only; file transcription on both)\n- microphone capture: `MicrophoneFileRecorder` (macOS only; the iOS build throws)\n- ASR: `EngineManager`\n- TTS generation: `TTSEngineManager`, `SynthesisRequest`\n- optional Apple playback: `VoxAppleSpeech` / `AppleSpeechOutputController`\n- outputs: `TranscriptionOutput`, `SynthesisOutput`, `TTSVoiceInfo`\n- telemetry: `PerformanceRecorder`, `PerformanceSample`\n- provider composition: `ProviderRegistry`, `TTSProviderRegistry`, `ProvidersConfig`, `ProviderEntry`\n\nFor anything beyond `VoxDictation`, keep raw Vox types behind one app-local actor such as `VoiceService`. Use `VoxAppleSpeech` when the app wants a reusable per-audible-surface playback controller; keep product policy in the app.\n\n## Package setup\n\n`VoxCore` and `VoxEngine` support direct transcription embedding on macOS 14+\nand iOS 17+. Minivox uses that same direct path and requires macOS 26+.\n\nAdd the package from GitHub:\n\n```swift\n.package(url: \"https://github.com/hudsonkit/vox.git\", from: \"0.5.2\")\n```\n\nFor a sibling checkout during local development, use a path dependency instead:\n\n```swift\n.package(path: \"../vox/swift\")\n```\n\nAdd these product dependencies to the app target:\n\n- `VoxCore`\n- `VoxEngine`\n- `VoxAppleSpeech` when the app wants the optional reusable Apple speech-output controller\n\nOnly add `VoxService` or `VoxBridge` if the app intentionally embeds companion/runtime behavior. That is not the default Apple app path.\n\n## Minimal local-first service\n\n```swift\nimport Foundation\nimport VoxCore\nimport VoxEngine\n\nactor VoiceStack {\n private let clientId: String\n private let asr: EngineManager\n private let tts: TTSEngineManager\n private let performance = PerformanceRecorder()\n\n init(clientId: String = \"my-app\") {\n self.clientId = clientId\n self.asr = EngineManager() // Parakeet\n self.tts = TTSEngineManager() // every built-in TTS provider; no automatic fallback\n }\n\n func warmup() async throws {\n _ = try await asr.preload(modelId: \"parakeet:v3\") { _ in }\n _ = try await tts.preload(modelId: TTSDefaults.modelId, voiceId: nil) { _ in }\n }\n\n func transcribe(fileURL: URL) async throws -> TranscriptionOutput {\n let output = try await asr.transcribe(url: fileURL, modelId: \"parakeet:v3\")\n\n await performance.record(\n PerformanceSample(\n clientId: clientId,\n route: \"transcribe.file\",\n modelId: output.modelId,\n outcome: \"ok\",\n textLength: output.text.count,\n metrics: output.metrics.performanceMetrics\n )\n )\n\n return output\n }\n\n func synthesize(text: String, voiceId: String? = nil) async throws -> SynthesisOutput {\n let output = try await tts.synthesize(\n SynthesisRequest(\n text: text,\n modelId: TTSDefaults.modelId,\n voiceId: voiceId\n )\n )\n\n await performance.record(\n PerformanceSample(\n clientId: clientId,\n route: \"synthesize.generate\",\n modelId: output.modelId,\n voiceId: output.voiceId,\n outcome: \"ok\",\n textLength: text.count,\n metrics: output.metrics.performanceMetrics\n )\n )\n\n return output\n }\n}\n```\n\n## App responsibilities\n\nIn embed mode, the app still owns:\n\n- microphone permission (`NSMicrophoneUsageDescription`, and the audio-input entitlement when sandboxed)\n- audio capture on iOS; on macOS `VoxDictation` or `MicrophoneFileRecorder` can capture for you\n- product-level spoken-output policy\n- interruption handling\n- product-level state and UX\n\nThe ASR entrypoint takes a `URL`, not an in-memory audio buffer. `VoxDictation` records to a temporary file and deletes it after `stop()`. On iOS, write the recording to a file yourself, then call `transcribe(fileURL:)`.\n\n`TTSProvider` and `TTSEngineManager` stay generation-only. They return WAV bytes in `SynthesisOutput.audioData` and do not own playback. Apps may play those bytes themselves, or opt into `VoxAppleSpeech` for a reusable Apple playback controller.\n\n## Optional Apple speech output\n\n`VoxAppleSpeech` is an optional embed product. `AppleSpeechOutputController` is a per-audible-surface playback arbiter, not a VoiceService facade and not an app product-policy layer.\n\nUse it when an Apple app wants one controller per audible surface that:\n\n- composes `TTSEngineManager` for generated-audio models\n- uses live `AVSpeechSynthesizer.speak()` for `avspeech:system`, never `write()`-then-play\n- plays `SynthesisOutput.audioData` through an injectable `AVAudioPlayer`-style sink\n- replaces the previous pending generation or playback when a new request starts\n- cancels the in-flight `Task`, audio player, and system synthesizer on idempotent `stop()` / `cancel()`, including during the enqueue window\n- reports typed phases (`resolving` or `generating`, `starting`, `playing`, `finished`, `cancelled`, `failed`)\n- reports requested vs actual model/voice identity separately from the physical audio-output route; it does not invent a provider id\n\nDo not treat this controller as browser playback, OpenScout Ranger policy, or app product policy. Reply dedupe, markdown flattening, fallback copy, preference storage, queue priority, and telemetry double-recording stay in the app.\n\nAudio-session configuration is injectable and disabled by default. There is no singleton and no process-global playback mutex; two controllers may run independently.\n\nRoute/model capability (`SpeechOutputCapabilities`) tells the controller whether a model is live system delivery or generated bytes. That capability does not add `speak()` to `TTSProvider`.\n\n## Warm-up\n\nWarm-up must remain explicit.\n\n- dictation: `VoxDictation.warmUp(progress:)`\n- ASR warm-up: `EngineManager.preload(modelId:progress:)`\n- TTS warm-up: `TTSEngineManager.preload(modelId:voiceId:progress:)`\n\nDo not hide warm-up behind app launch side effects. Warm on intent at a predictable app state transition, such as opening the screen that dictates. Without warm-up, the first transcription pays the model load.\n\n## Telemetry\n\nCompanion mode records telemetry automatically. In embed mode, `VoxDictation` records a sample for every transcription (pass `recordsPerformance: false` to opt out). The lower-level types do not.\n\nWhen you call `EngineManager` or `TTSEngineManager` directly, record samples yourself with `PerformanceRecorder` and preserve these dimensions:\n\n- `clientId`\n- `route`\n- `modelId`\n- `voiceId` for synthesis\n\nUse the same route names Vox Companion uses:\n\n- `transcribe.dictation` (microphone dictation, as `VoxDictation` records it)\n- `transcribe.file`\n- `synthesize.generate`\n\n`RuntimePaths.performanceLogURL()` resolves to:\n\n- macOS: `~/.vox/performance.jsonl`\n- iOS: `Application Support/Vox/performance.jsonl`\n\n## OpenAI TTS in embed mode\n\n- ASR: `EngineManager()` -> `ParakeetProvider()`\n- TTS: `TTSEngineManager()` registers every built-in provider: OpenAI, ElevenLabs, MiniMax, NVIDIA, Groq, Gemini and AVSpeech. It routes by model id and never falls back from one provider to another.\n- default TTS model: `TTSDefaults.modelId` = `gpt-4o-mini-tts`, which needs an OpenAI key\n- local TTS model: `TTSDefaults.localModelId` = `avspeech:system`, on device, no key\n\nA synthesis request with `gpt-4o-mini-tts` and no key throws. If the app wants the system voice as a fallback, catch the error and retry with `TTSDefaults.localModelId`.\n\nKeys are looked up in this order: `SynthesisRequest.providerCredentials`, the provider entry's `env`, the process environment, then the Vox credential store. An iOS app has no useful process environment, so pass the key in code. Never ship a long-lived key inside the app binary.\n\n```swift\nlet ttsConfig = ProvidersConfig(providers: [\n ProviderEntry(\n id: \"avspeech\",\n kind: .tts,\n builtin: true,\n models: [AVSpeechSynthesizerProvider.modelID]\n ),\n ProviderEntry(\n id: \"openai-tts\",\n kind: .tts,\n builtin: true,\n models: OpenAITTSProvider.supportedModelIDs,\n env: [\"OPENAI_API_KEY\": apiKey]\n )\n])\n\nlet tts = TTSEngineManager(provider: TTSProviderRegistry(config: ttsConfig))\n```\n\n## Default plan\n\nFor a new Apple app integration, the default plan is:\n\n- use embed mode on iOS and macOS\n- add `VoxCore` and `VoxEngine`, plus `VoxAppleSpeech` if the app wants reusable Apple playback\n- start with `VoxDictation`; wrap anything lower-level in one app-local actor or service\n- use `parakeet:v3` for ASR; `parakeet:v2` is the English-only TDT option\n- use `gpt-4o-mini-tts` for default TTS\n- use `avspeech:system` only when a local system voice fallback is required\n- keep `VoxDictation`'s timings on, or record Vox-compatible samples yourself\n- use Vox Companion only for web surfaces or cross-process workflows\n\n## What agents should not assume\n\n- `VoxDictation` covers dictation only. `VoxAppleSpeech` is optional playback, separate from it.\n- `VoxDictation.start()` is macOS only; on iOS it throws.\n- There is no public embed live-session coordinator yet: no partial text while recording.\n- There is no public embed warm-up coordinator helper.\n- `EngineManager` and `TTSEngineManager` do not write performance samples; only `VoxDictation` does.\n- Apple apps do not need `@voxd/sdk` or `@voxd/client`.\n- `VoxAppleSpeech` does not own browser playback or app product policy.\n\n## First tasks in a sibling app repo\n\n1. Add `https://github.com/hudsonkit/vox`, or `../vox/swift` for a sibling checkout.\n2. Create one `VoxDictation`, or a single `VoiceService` actor in app code for lower-level use.\n3. Warm explicitly, on user intent.\n4. Feed ASR with file URLs, not raw buffers.\n5. Feed TTS output WAV data into the app playback layer, or use `AppleSpeechOutputController` for per-surface Apple playback.\n6. Emit `PerformanceSample` records with stable route names.\n7. Keep Companion mode out of the Apple app path unless the feature is genuinely web or cross-process.\n\nSee [Observability](./observability.md) for metric interpretation and [Architecture](./architecture.md) for package ownership.\n\n## Packaged app resources\n\nXcode copies SwiftPM resource bundles into the app for you. If you package the app yourself, copy the bundle SwiftPM built into the signed app's `Contents/Resources`. SwiftPM names it after the package that declares the target: `Vox_HudsonSpeechEngine.bundle` when you depend on the repository root, `HudsonSpeechEngine_HudsonSpeechEngine.bundle` when you depend on `vox/swift`. Vox looks for both.\n\n`SpeechEngineResources` resolves the model catalog and mlx-audio provider script\nfrom that location in apps, and from SwiftPM resources in command-line builds.\nA packaged app does not fall back to a developer build directory."
},
{
"id": "runtime",
"title": "Runtime",
"description": "Companion request flow, warm-up, sessions, routes, ports, and runtime ownership boundaries.",
"content": "## Core Flow\n\n1. A companion-connected client connects to `voxd` over local WebSocket JSON-RPC, or reaches it indirectly through the local HTTP bridge.\n2. The runtime resolves health, model state, and optional warm-up state.\n3. The client triggers file transcription, file annotation, a live ASR session, one-shot synthesis, or a synthesis session.\n4. `VoxEngine` resolves the right ASR, annotation, or TTS provider and returns transcript text, speaker attribution, word timings, or WAV bytes plus stage metrics.\n5. The runtime records a tagged performance sample to `~/.vox/performance.jsonl`.\n6. The daemon appends operational logs to `~/.vox/logs/voxd.log`.\n\n## Warm-Up\n\nWarm-up is a public API, not a hidden side effect. It applies to both ASR and TTS models.\n\n- `warmup.status` -- check if the model is hot\n- `warmup.start` -- warm immediately\n- `warmup.schedule` -- warm after a delay\n\nTypical pattern: create a `VoxClient` with a stable `clientId`, warm when the user opens a voice affordance, then transcribe or synthesize once the model is ready.\n\nThe ASR model itself is never embedded in the Swift product. Automatic\nacquisition for live transcription is controlled by\n`VoxSpeechPreferences.modelDownloadPolicy`:\n\n| Value | Behavior |\n|-------|----------|\n| `never` | Never download automatically. Live transcription requires an already-installed model or an explicit install/warm-up request. |\n| `on_first_use` | Start warm-up with the first live transcription session. This is the default. |\n| `eager` | Start warm-up when `voxd` starts. |\n\nThe explicit `models.install`, `models.preload`, and `warmup.*` routes remain\navailable under every policy; choosing one of those routes is itself an\noperator/user request to acquire the model.\n\n## File Transcription\n\n`transcribe.file` is best for benchmarks because it takes mic capture out of the measurement. Returns transcript text, word-level timestamps, `modelId`, elapsed time, and stage metrics.\n\n## File Annotation\n\n`annotate.file` is the file-first speaker annotation route. It accepts an audio path plus optional transcript text and word timings, and returns speaker segments, speaker-attributed words, and annotation metrics.\n\nThe route is scaffolded now so evaluation harnesses and future diarization backends can share one contract even before the default annotation backend lands.\n\n## Synthesis\n\n`synthesize.voices` lists available voices for a TTS model.\n\n`synthesize.generate` returns:\n\n- `modelId`\n- `voiceId`\n- `format`\n- `contentType`\n- base64-encoded audio bytes\n- elapsed time\n- synthesis metrics\n\nLonger-running output flows use:\n\n- `synthesize.startSession`\n- `synthesize.sessionStatus`\n- `synthesize.cancel`\n\n## RPC Routes\n\nThe runtime exposes these RPC routes:\n\n- `transcribe.file`\n- `annotate.file`\n- `transcribe.startSession`\n- `transcribe.sessionStatus`\n- `transcribe.stopSession`\n- `transcribe.cancelSession`\n- `synthesize.voices`\n- `synthesize.generate`\n- `synthesize.startSession`\n- `synthesize.sessionStatus`\n- `synthesize.cancel`\n- `models.list`\n- `models.install`\n- `models.preload`\n- `models.catalog`\n- `models.refreshCatalog`\n- `warmup.status`\n- `warmup.start`\n- `warmup.schedule`\n\n`models.catalog` returns the published dictation catalog, including `plugins[]`. Plugin install is `vox plugins install`, not an RPC. `voxd` loads `~/.vox/plugins/*/provider.json` at start. See [Models and plugins](./models.md).\n\n## Performance routes\n\nThese are the route values currently emitted into `performance.jsonl`:\n\n- `transcribe.file`\n- `annotate.file`\n- `transcribe.live`\n- `synthesize.generate`\n- `synthesize.startSession`\n\nCancelled synthesis sessions are currently recorded as `synthesize.startSession` with `outcome: \"cancelled\"`.\n\n## Live Sessions\n\nSpeech sessions are coordinated in `VoxService`. Session ownership ties to both `connectionID` and `clientId`. Stop and cancel are distinct operations. Final transcript events include metrics and word-level timestamps. Synthesis session status includes the selected model and voice. Active session state is inspectable for operator recovery.\n\n## Configuration\n\nPorts and bind address are configurable via environment variables.\n\nVox uses two named companion ports:\n\n| Role | Default | Transport | Env var | Stored |\n|------|---------|-----------|---------|--------|\n| `companion-ws` | `42137` | WebSocket | `VOX_PORT` | `runtime.json` |\n| `companion-http` | `43115` | HTTP | `VOX_BRIDGE_PORT` | process/env only |\n\n`companion-ws` is the `voxd` daemon port that `@voxd/sdk` discovers through `~/.vox/runtime.json`. `companion-http` is the browser-facing bridge port used by `@voxd/client`.\n\nThe bind host is shared across both surfaces:\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `VOX_PORT` | `42137` | `companion-ws` daemon WebSocket port |\n| `VOX_BRIDGE_PORT` | `43115` | `companion-http` bridge port |\n| `VOX_HOST` | `127.0.0.1` | Bind address for both services |\n| `VOX_HOME` | `~/.vox` | Runtime data directory |\n\nCLI flag `--port` takes precedence over env vars for both `voxd` and `voxbridge`.\n\n## Important Swift entry points\n\n- `swift/Sources/voxd/main.swift`\n- `swift/Sources/VoxService/VoxRuntimeService.swift`\n- `swift/Sources/VoxService/LiveSessionCoordinator.swift`\n- `swift/Sources/VoxService/SynthesisSessionCoordinator.swift`\n- `swift/Sources/VoxService/WarmupCoordinator.swift`\n- `swift/Sources/HudsonSpeechEngine/ProviderRegistry.swift`\n- `swift/Sources/HudsonSpeechEngine/TTSProviderRegistry.swift`\n- `swift/Sources/VoxCore/SpeechModelCatalog.swift`\n- `swift/Sources/VoxCore/PluginStore.swift`\n\nUse the operator surface to inspect the runtime without guessing:\n\n```bash\nvox doctor\nvox warmup status parakeet:v3\nvox transcribe status\nvox logs daemon --tail 80\n```\n\nSee [Architecture](./architecture.md) for layer ownership and [Observability](./observability.md) for performance sample interpretation."
},
{
"id": "models",
"title": "Models and plugins",
"description": "Published dictation catalog, built-in families, and how Vox plugins install new ASR runtimes.",
"content": "The published catalog at `https://voxd.cc/data/models.json` is the source of dictation model ids. The same file lives in the repo as `data/models.json` and is bundled with the engine as a fallback.\n\nRefresh the local snapshot explicitly:\n\n```bash\nvox models catalog\nvox models catalog refresh\n```\n\n`VOX_MODEL_CATALOG_URL` overrides the catalog URL. A failed refresh keeps the bundled or cached snapshot. Catalog refresh never installs a plugin and never runs a plugin command.\n\n## Built-in families\n\nThese families run without a plugin:\n\n| Family | Default or typical ids | How it runs |\n|---|---|---|\n| `parakeet-tdt` | `parakeet:v3` (default), `parakeet:v2` (English) | On-device CoreML |\n| `apple-speech` | `apple:speech-transcriber` | On-device Speech framework; macOS 26+ |\n| `openai-transcribe` | `gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `whisper-1` | OpenAI Audio API; needs `OPENAI_API_KEY` |\n| `mlx-audio` | Qwen3-ASR, Cohere Transcribe, Nemotron 3.5 ASR Streaming, Whisper turbo, Parakeet MLX ids | mlx-audio provider; opt-in in `providers.json` |\n\n`parakeet:v3` stays the default for Minivox, Vox.app, and CLI warmup. Use `parakeet:v2` when you want the English-only TDT bundle.\n\nApple SpeechTranscriber is registered by the default daemon, so selecting its model id is enough. Apple downloads locale assets through `AssetInventory`; `VOX_APPLE_SPEECH_LOCALE` selects the locale.\n\n## Current local shortlist\n\n- `parakeet:v3`: default, mature CoreML path, 25 European languages.\n- `apple:speech-transcriber`: system-managed on-device parity option with word timings on macOS 26+.\n- `mlx-community/Qwen3-ASR-1.7B-8bit`: multilingual accuracy candidate on Apple Silicon.\n- `mlx-community/cohere-transcribe-03-2026-mlx-8bit`: strong offline multilingual candidate; not a realtime model.\n- `mlx-community/nemotron-3.5-asr-streaming-0.6b-8bit`: 0.6B cache-aware streaming RNNT candidate covering 35 languages.\n\nCatalog `capabilities.liveTranscription` describes what Vox exposes today, not what an upstream model architecture can theoretically do. It remains `false` for every current provider because Vox's public ASR contract still accepts a completed audio file and returns finalized text. Apple and Nemotron are streaming-capable foundations for the next runtime slice, but selecting them does not yet turn `session.partial` into model-backed partial transcription.\n\nVox does not ship Intel support. Local runtime entries are marked `architectures: [\"arm64\"]`. Audio conversion for native model inputs uses platform or provider libraries; do not add hand-written sample-rate conversion or denoising code to a model adapter.\n\nSDK catalog methods:\n\n```ts\nawait client.listCatalog();\nawait client.refreshCatalog();\n```\n\nRPC routes: `models.catalog`, `models.refreshCatalog`.\n\n## Plugins\n\nA plugin is an external JSON-RPC provider advertised in the catalog `plugins[]` array. Models point at a plugin with a `plugin` field.\n\nInstall is a CLI filesystem action, not an RPC:\n\n```bash\nvox plugins list\nvox plugins install mlx-vlm\n# restart voxd\nvox transcribe file --model gemma-4-e2b-it /path/to/audio.wav\nvox plugins remove mlx-vlm\n```\n\n`vox plugins install` writes `~/.vox/plugins/<id>/provider.json`. `voxd` loads those files at start and merges them into the provider registry.\n\nGemma 4 E2B (`gemma-4-e2b-it`) uses plugin `mlx-vlm`. The bundled runner speaks the provider protocol and calls mlx-vlm when `VOX_MLX_VLM_PYTHON` or `python3` has that package. Without mlx-vlm the plugin stays installed and reports `available=false`.\n\nPlugin launchers are allowlisted: `node`, `bun`, `npx`, `bunx`, `uv`, `uvx`, `python3`, `python`. Native bundles can also launch a compiled executable inside their own plugin directory. Bundled plugins take their command from the CLI's bundle manifest. Installation writes an absolute command path so the daemon can start it without a shell.\n\n### Cactus Whistle\n\nWhistle runs in an optional Rust provider process with Cactus's prebuilt Needle CPU runtime on Apple Silicon. The compiled provider ships with the CLI and needs no Python, `uv`, or Rust installation. It loads only the speech model:\n\n```bash\nvox plugins install whistle\n# restart voxd\nvox models install whistle\nvox models preload whistle\nvox transcribe file --model whistle --metrics --timestamps /path/to/audio.wav\n```\n\nPlugin registration copies the compiled provider and writes `provider.json`. `models install` downloads the speech weights and native runtime. `models preload` loads the installed model into the persistent helper. Model listing does not download weights or load the model.\n\nAssets are stored under `VOX_HOME/models/whistle` (default `~/.vox/models/whistle`). Transcription can load installed weights on its first call; `models preload` moves that work out of the transcription request. Neither operation downloads missing model assets.\n\nSet `VOX_WHISTLE_LANGUAGE` in the provider's `env` to `en`, `de`, `fr`, `es`, `it`, `nl`, or `pl` to select a language. An unset value uses language detection. `VOX_WHISTLE_KEYWORDS` accepts comma-separated words for keyword biasing. Restart `voxd` after changing provider environment values.\n\nThe adapter accepts audio files up to 30 seconds and returns finalized text with word timestamps and confidence. Audio decoding and 16 kHz mono conversion use Symphonia and Rubato. This provider exposes file transcription; its catalog entry keeps `liveTranscription=false`.\n\nWhistle's runtime has one shared model state per process. The helper accepts one operation at a time and returns a busy error for overlapping requests. After an error, callers can retry on the same process. Removing the plugin unregisters the helper; downloaded model assets remain available for reinstallation.\n\nThe [Whistle model card](https://huggingface.co/Cactus-Compute/whistle) documents the supported languages and model. The [Needle repository](https://github.com/cactus-compute/needle) documents the native runtime; the Rust adapter calls its C interface directly. Model installation downloads the pinned weights and runtime binary, checks their hashes, and keeps them for offline inference.\n\nSee the [Provider Protocol](./providers.md) for the stdin/stdout contract a plugin must implement."
},
{
"id": "providers",
"title": "Provider Protocol",
"description": "How external STT and TTS engines plug into Vox via JSON-RPC over stdin/stdout.",
"content": "Vox separates the runtime (mic capture, sessions, routing, telemetry, playback handoff) from the speech engine. Engines are called _providers_. They can be external processes or built-in bridges that speak JSON-RPC over stdin/stdout.\n\nProvider configuration plugs directly into the companion runtime's install, preload, and route-dispatch flow. Read [Runtime](./runtime.md) alongside this spec if you want the full daemon-side picture.\n\nProviders can serve either:\n\n- ASR / STT: accept audio and return text\n- TTS: accept text and return audio\n\nBuilt-in providers include:\n\n- `parakeet` for on-device CoreML ASR (`parakeet:v3` default, `parakeet:v2` English)\n- `apple-speech` for Apple SpeechTranscriber (`apple:speech-transcriber`, macOS 26+ on Apple Silicon)\n- `openai-transcribe` for remote OpenAI file transcription (`gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `whisper-1`)\n- `avspeech` for system TTS\n- `openai-tts` for remote TTS\n- `elevenlabs` for ElevenLabs remote TTS\n- `minimax` for MiniMax remote TTS\n- `nvidia` for NVIDIA Magpie remote TTS\n- `groq` for Groq Orpheus remote TTS\n- `gemini` for Gemini remote TTS\n- `mlx-audio` for built-in external bridging across both ASR and TTS, including Whisper, Qwen3-ASR, Cohere Transcribe, Nemotron 3.5, and Parakeet MLX ids\n\n## Model catalog\n\nOperator-facing catalog and plugin steps live in [Models and plugins](./models.md).\n\nThe published catalog at `https://voxd.cc/data/models.json` is the source of dictation model ids. The same file lives in the repo as `data/models.json` and is bundled with the engine as a fallback.\n\nEach model names a `family`. Built-in families the engine can run without a plugin:\n\n- `parakeet-tdt`: CoreML Parakeet TDT bundles\n- `apple-speech`: Apple SpeechTranscriber and system-managed locale assets\n- `mlx-audio`: Hugging Face ids loaded by the mlx-audio provider\n- `openai-transcribe`: OpenAI Audio transcriptions API\n\nNew families ship as plugins in the same catalog. A plugin is an external JSON-RPC provider. The catalog advertises it; install is explicit; `voxd` loads `~/.vox/plugins/<id>/provider.json` on start.\n\n```bash\nvox plugins list\nvox plugins install mlx-vlm\n# restart voxd\nvox transcribe file --model gemma-4-e2b-it /path/to/audio.wav\nvox plugins remove mlx-vlm\n```\n\nRefresh is explicit:\n\n```bash\nvox models catalog\nvox models catalog refresh\n```\n\n`VOX_MODEL_CATALOG_URL` overrides the catalog URL. A failed refresh keeps the bundled or cached snapshot. Catalog refresh never installs or runs a plugin command. `vox plugins install` copies a bundled runner or records an allowlisted launcher (`node`, `bun`, `npx`, `bunx`, `uv`, `uvx`, `python3`, `python`).\n\n## Provider Config\n\nProviders are registered in `~/.vox/providers.json`:\n\n```json\n{\n \"providers\": [\n {\n \"id\": \"parakeet\",\n \"kind\": \"asr\",\n \"builtin\": true,\n \"models\": [\"parakeet:v3\", \"parakeet:v2\"]\n },\n {\n \"id\": \"apple-speech\",\n \"kind\": \"asr\",\n \"builtin\": true,\n \"models\": [\"apple:speech-transcriber\"],\n \"env\": {\n \"VOX_APPLE_SPEECH_LOCALE\": \"en-US\"\n }\n },\n {\n \"id\": \"avspeech\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"avspeech:system\"]\n },\n {\n \"id\": \"openai-tts\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"gpt-4o-mini-tts\"],\n \"env\": {\n \"OPENAI_API_KEY\": \"sk-...\",\n \"VOX_OPENAI_TTS_TIMEOUT_SECONDS\": \"12\"\n }\n },\n {\n \"id\": \"elevenlabs\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"eleven_multilingual_v2\"],\n \"env\": {\n \"ELEVENLABS_API_KEY\": \"...\"\n }\n },\n {\n \"id\": \"minimax\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"speech-2.8-hd\"],\n \"env\": {\n \"MINIMAX_API_KEY\": \"...\"\n }\n },\n {\n \"id\": \"nvidia\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"magpie-tts-multilingual\"],\n \"env\": {\n \"NV_API_KEY\": \"...\"\n }\n },\n {\n \"id\": \"groq\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"canopylabs/orpheus-v1-english\"],\n \"env\": {\n \"GROQ_API_KEY\": \"...\"\n }\n },\n {\n \"id\": \"gemini\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"models\": [\"gemini-2.5-flash-preview-tts\"],\n \"env\": {\n \"GEMINI_API_KEY\": \"...\"\n }\n },\n {\n \"id\": \"mlx-audio\",\n \"kind\": \"asr\",\n \"builtin\": true,\n \"env\": {\n \"VOX_MLX_AUDIO_PYTHON\": \"/path/to/venv/bin/python\",\n \"VOX_MLX_AUDIO_ASR_MODELS\": \"mlx-community/Qwen3-ASR-1.7B-8bit,mlx-community/cohere-transcribe-03-2026-mlx-8bit,mlx-community/nemotron-3.5-asr-streaming-0.6b-8bit\"\n }\n },\n {\n \"id\": \"mlx-audio\",\n \"kind\": \"tts\",\n \"builtin\": true,\n \"env\": {\n \"VOX_MLX_AUDIO_PYTHON\": \"/path/to/venv/bin/python\",\n \"VOX_MLX_AUDIO_TTS_MODELS\": \"mlx-community/Soprano-1.1-80M-bf16,mlx-community/Kokoro-82M-4bit\",\n \"VOX_MLX_AUDIO_TTS_DEFAULT_VOICE\": \"af_heart\"\n }\n }\n ]\n}\n```\n\n| Field | Type | Required | Description |\n|-------|------|----------|-------------|\n| `id` | `string` | Yes | Unique identifier for this provider. |\n| `kind` | `\"asr\" \\| \"tts\"` | No | Provider kind. Defaults to `asr` if omitted. |\n| `builtin` | `boolean` | No | If `true`, Vox uses its bundled implementation for the given `id`. |\n| `command` | `string[]` | No | Executable and arguments Vox will spawn for an external provider. |\n| `models` | `string[]` | No | Model IDs this provider serves. Optional when the provider reports models dynamically. |\n| `env` | `Record<string, string>` | No | Extra environment variables passed to the provider process. |\n\nNotes:\n\n- Vox supports Apple Silicon only. Do not treat successful remote-provider compilation on an Intel host as a supported configuration.\n- Register ASR and TTS as separate entries even when they share the same `id`.\n- `models` is optional for external providers now. Vox can call `models()` and route dynamically from the returned list.\n- OpenAI, NVIDIA Magpie, Groq, and Gemini/Google accept per-request `credentials` on `synthesize.generate` / companion synthesis. Those keys are allowlisted through `VoxRuntimeService` and `VoxBridge`; unknown keys, including ElevenLabs and MiniMax keys, are dropped. ElevenLabs and MiniMax currently read only provider `env` and the process environment.\n- Credential order for the providers that accept lent keys is per-request `credentials`, then provider `env`, then the process environment. An explicit provider `env` key or alias, even empty or whitespace, fences process fallback for both configured credential resolution and default model selection so a host can lend keys per request without inheriting an ambient secret. Process env is consulted only when no relevant credential key is present in provider `env` (including when `env` is omitted). Values are never logged or stored by the allowlist parser.\n- NVIDIA Magpie, Groq, and Gemini synthesis HTTP errors redact the exact request text before a vendor body can be persisted or logged. JSON diagnostics keep useful vendor detail after secret and exact-prompt redaction. Plain-text non-JSON bodies fail closed whenever synthesis sensitive values (prompt or instructions) are supplied, including distinctive fragments such as `quota: ` plus a short prompt prefix. Voice discovery has no prompt and is not covered by that fail-closed path. This sanitization applies to these new remote providers; OpenAI, ElevenLabs, MiniMax, and AVSpeech keep their existing error surfaces.\n- `ELEVENLABS_BASE_URL`, `ELEVENLABS_OUTPUT_FORMAT`, `MINIMAX_BASE_URL`, `NVIDIA_TTS_URL`, `NVIDIA_VOICES_URL`, `GROQ_BASE_URL`, and `GEMINI_BASE_URL` can override vendor defaults.\n- NVIDIA Magpie accepts `NV_API_KEY` and the `NVIDIA_API_KEY` compatibility alias, plus camelCase / snake_case variants. Gemini accepts `GEMINI_API_KEY`, `GOOGLE_API_KEY`, and `GOOGLE_GENAI_API_KEY`.\n- Remote providers are considered for default model selection and in-process default registration only when they are configured in `providers.json` or the daemon environment. An API key in the process environment is not broader user consent than selecting or configuring that provider. Configured aliases such as `magpie`, `groq-tts`, and `google-tts` prevent the canonical default entries from being appended over that family's routing and key config.\n- If `providers.json` contains only ASR entries, Vox falls back to default TTS providers. The inverse is also true. The default TTS model remains `gpt-4o-mini-tts` when OpenAI is configured. App-registry and daemon/config default selection share one canonical ranking independent of `providers.json` order and independent of the alphabetically sorted `TTSProviderRegistry` model list: OpenAI `gpt-4o-mini-tts`, then ElevenLabs, MiniMax, NVIDIA Magpie, Groq, Gemini, then `avspeech:system`.\n- The public ASR provider method is currently file-based. Apple SpeechTranscriber and Nemotron have streaming-capable internals, but live partial events require a separate runtime/protocol integration.\n- Native audio adaptation uses `AVAudioConverter`, SpeechAnalyzer's file path, or provider-owned conversion. Provider adapters must not implement ad hoc resampling or denoising.\n\n### OpenAI TTS timeout\n\n`openai-tts` uses a hard wall-clock request timeout so stalled remote TTS calls do not block the caller for minutes.\n\n- default: `12` seconds\n- env override: `VOX_OPENAI_TTS_TIMEOUT_SECONDS`\n- compatibility alias: `OPENAI_TTS_TIMEOUT_SECONDS`\n- maximum accepted value: `30` seconds\n\nThe timeout can be set in the provider `env` block or the daemon process environment.\n\n### NVIDIA Magpie TTS\n\n`nvidia` is the first-class Vox home for NVIDIA Magpie TTS Multilingual, using the hosted NVIDIA Developer Inference contract:\n\n- model id: `magpie-tts-multilingual`\n- voice discovery: `GET /v1/audio/list_voices`\n- synthesis: multipart `POST /v1/audio/synthesize`\n- encoding: `LINEAR_PCM` at `44100` Hz\n- default voice: `Magpie-Multilingual.EN-US.Aria`\n- language is inferred from the Magpie voice id (`EN-US` → `en-US`)\n- credentials: `NV_API_KEY`, with `NVIDIA_API_KEY` as a compatibility alias; per-request lent keys are accepted\n- input limit: 2000 normalized characters; longer text is rejected rather than truncated\n- URL overrides: `NVIDIA_TTS_URL`, `NVIDIA_VOICES_URL`, or `NVIDIA_BASE_URL`\n\nIf Magpie returns a valid WAVE container, Vox passes it through. If it returns aligned LINEAR PCM (`audio/l16` or similar), Vox wraps that PCM as 16-bit 44.1 kHz WAV so the public `SynthesisOutput` contract stays `format: \"wav\"`. Other payloads are rejected.\n\n### Groq Orpheus TTS\n\n`groq` uses Groq's OpenAI-compatible speech endpoint and requests `response_format: \"wav\"`:\n\n- models: `canopylabs/orpheus-v1-english`, `canopylabs/orpheus-arabic-saudi`\n- default voice: `autumn`\n- input limit: 200 characters; longer text is rejected rather than truncated\n- credentials: `GROQ_API_KEY`; per-request lent keys are accepted\n- URL override: `GROQ_BASE_URL` (default `https://api.groq.com/openai/v1`)\n- successful responses must be a structurally valid WAV; nonempty non-WAV bytes are rejected\n\n### Gemini TTS\n\n`gemini` uses the Gemini Generate Content API with `responseModalities: [\"AUDIO\"]`:\n\n- models: `gemini-2.5-flash-preview-tts`, `gemini-2.5-pro-preview-tts`, `gemini-3.1-flash-tts-preview`\n- default voice: `Puck`\n- official prebuilt catalog includes `Rasalgethi` (not `Rasalas`)\n- credentials: `GEMINI_API_KEY`, with `GOOGLE_API_KEY` / `GOOGLE_GENAI_API_KEY` as aliases; per-request lent keys are accepted\n- URL override: `GEMINI_BASE_URL` (default `https://generativelanguage.googleapis.com/v1beta`); trailing slashes are normalized before joining `/models/{model}:generateContent`\n\nGemini 2.5 and Gemini 3.1 Flash TTS Preview share the Generate Content `responseModalities: [\"AUDIO\"]` contract. Vox does not use a separate Interactions API for these models.\n\nGemini typically returns `audio/L16` PCM. Vox wraps that PCM in a WAV container so callers still receive `SynthesisOutput.format = \"wav\"`. This is a lossless PCM container wrap, not a transcode.\n\n## Protocol Methods\n\nAll communication uses newline-delimited JSON-RPC 2.0 over stdin (requests from Vox) and stdout (responses from the provider).\n\n### `models`\n\nList available models for the provider kind.\n\n**Request:**\n\n```json\n{ \"jsonrpc\": \"2.0\", \"id\": 1, \"method\": \"models\" }\n```\n\n**Response:**\n\n```json\n{\n \"jsonrpc\": \"2.0\",\n \"id\": 1,\n \"result\": {\n \"models\": [\n {\n \"id\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n \"name\": \"whisper-large-v3-turbo-asr-fp16\",\n \"backend\": \"mlx-audio\",\n \"installed\": true,\n \"preloaded\": false,\n \"available\": true\n }\n ]\n }\n}\n```\n\n### `install`\n\nDownload or prepare model files.\n\n**Request:**\n\n```json\n{ \"jsonrpc\": \"2.0\", \"id\": 2, \"method\": \"install\", \"params\": { \"modelId\": \"mlx-community/Kokoro-82M-4bit\" } }\n```\n\nThe provider can emit progress notifications on stdout during installation or preload:\n\n```json\n{ \"jsonrpc\": \"2.0\", \"method\": \"progress\", \"params\": { \"modelId\": \"mlx-community/Kokoro-82M-4bit\", \"progress\": 0.5, \"status\": \"loading\" } }\n```\n\n**Response:** `{ \"model\": { ... } }`, where `model` matches an entry returned by `models`.\n\n### `preload`\n\nLoad a model into memory so subsequent requests start faster.\n\n**Request:**\n\n```json\n{ \"jsonrpc\": \"2.0\", \"id\": 3, \"method\": \"preload\", \"params\": { \"modelId\": \"mlx-community/Soprano-1.1-80M-bf16\" } }\n```\n\n**Response:** `{ \"model\": { ... } }`, with `model.preloaded: true` after loading succeeds.\n\n## ASR Methods\n\n### `transcribe`\n\nTranscribe an audio file.\n\n**Request:**\n\n```json\n{ \"jsonrpc\": \"2.0\", \"id\": 4, \"method\": \"transcribe\", \"params\": { \"modelId\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\", \"path\": \"/tmp/audio.wav\" } }\n```\n\n**Response:**\n\n```json\n{\n \"jsonrpc\": \"2.0\",\n \"id\": 4,\n \"result\": {\n \"modelId\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n \"text\": \"Hello world\",\n \"elapsedMs\": 142,\n \"metrics\": {\n \"inferenceMs\": 130,\n \"modelLoadMs\": 0,\n \"audioLoadMs\": 5,\n \"audioPrepareMs\": 2,\n \"fileCheckMs\": 1,\n \"modelCheckMs\": 1,\n \"totalMs\": 142\n },\n \"words\": [\n { \"word\": \"Hello\", \"start\": 0.12, \"end\": 0.44, \"confidence\": 0.99 },\n { \"word\": \"world\", \"start\": 0.45, \"end\": 0.71, \"confidence\": 0.98 }\n ]\n }\n}\n```\n\n## TTS Methods\n\n### `voices`\n\nList available voices for a model. If `modelId` is omitted, Vox may call `voices` across multiple models and merge the results.\n\n**Request:**\n\n```json\n{ \"jsonrpc\": \"2.0\", \"id\": 5, \"method\": \"voices\", \"params\": { \"modelId\": \"mlx-community/Kokoro-82M-4bit\" } }\n```\n\n**Response:**\n\n```json\n{\n \"jsonrpc\": \"2.0\",\n \"id\": 5,\n \"result\": {\n \"voices\": [\n {\n \"id\": \"af_heart\",\n \"name\": \"af_heart\",\n \"language\": \"en-US\",\n \"backend\": \"mlx-audio\",\n \"modelId\": \"mlx-community/Kokoro-82M-4bit\",\n \"available\": true,\n \"default\": true\n }\n ]\n }\n}\n```\n\n### `synthesize`\n\nGenerate audio from text.\n\n**Request:**\n\n```json\n{\n \"jsonrpc\": \"2.0\",\n \"id\": 6,\n \"method\": \"synthesize\",\n \"params\": {\n \"modelId\": \"mlx-community/Soprano-1.1-80M-bf16\",\n \"input\": \"Hello from Vox\",\n \"voiceId\": \"af_heart\",\n \"format\": \"wav\",\n \"speed\": 1.0\n }\n}\n```\n\n**Response:**\n\n```json\n{\n \"jsonrpc\": \"2.0\",\n \"id\": 6,\n \"result\": {\n \"modelId\": \"mlx-community/Soprano-1.1-80M-bf16\",\n \"voiceId\": \"af_heart\",\n \"format\": \"wav\",\n \"contentType\": \"audio/wav\",\n \"audioBase64\": \"<base64 wav data>\",\n \"elapsedMs\": 418,\n \"metrics\": {\n \"audioDurationMs\": 1024,\n \"characterCount\": 14,\n \"modelCheckMs\": 0,\n \"modelLoadMs\": 0,\n \"voiceResolveMs\": 1,\n \"synthesisMs\": 363,\n \"totalMs\": 418\n }\n }\n}\n```\n\n## Metrics Contract\n\nProviders must return stage timings in the `metrics` object of every `transcribe` or `synthesize` response. These feed into Vox telemetry tagged with `modelId`, `route`, and, for TTS, `voiceId`.\n\n### ASR metrics\n\nRequired fields:\n\n| Field | Type | Description |\n|-------|------|-------------|\n| `inferenceMs` | `number` | Time spent running the model. |\n| `totalMs` | `number` | Wall-clock time for the entire request. |\n\nOptional but recommended:\n\n| Field | Type | Description |\n|-------|------|-------------|\n| `modelLoadMs` | `number` | Time loading the model (0 if already preloaded). |\n| `audioLoadMs` | `number` | Time reading the audio file from disk. |\n| `audioPrepareMs` | `number` | Time resampling or converting the audio. |\n| `fileCheckMs` | `number` | Time validating the audio file exists and is readable. |\n| `modelCheckMs` | `number` | Time checking the model is installed and ready. |\n| `audioDurationMs` | `number` | Duration of the input audio. |\n\n### TTS metrics\n\nRequired fields:\n\n| Field | Type | Description |\n|-------|------|-------------|\n| `totalMs` | `number` | Wall-clock time for the entire request. |\n\nOptional but recommended:\n\n| Field | Type | Description |\n|-------|------|-------------|\n| `synthesisMs` | `number` | Time spent generating audio once the model is running. |\n| `inferenceMs` | `number` | Accepted as a fallback alias for `synthesisMs`. |\n| `modelLoadMs` | `number` | Time loading the model (0 if already preloaded). |\n| `modelCheckMs` | `number` | Time checking the model is installed and ready. |\n| `voiceResolveMs` | `number` | Time resolving the requested voice. |\n| `audioDurationMs` | `number` | Duration of the synthesized audio. |\n| `outputBytes` | `number` | Number of encoded output bytes. |\n| `characterCount` | `number` | Length of the input text. |\n\n## What Vox handles\n\nProviders only deal with models plus transcription or synthesis. The runtime handles everything else:\n\n- Mic permissions and capture -- ASR providers receive a WAV file path\n- Audio format normalization for ASR input\n- Playback handoff -- TTS providers return audio bytes and Vox hands them back to the caller\n- Session lifecycle -- start, stop, cancel coordinated by the daemon\n- Warm-up scheduling and state\n- Client identity routing (`clientId`)\n- Performance telemetry collection\n- Provider execution capacity and backpressure -- requests are not globally serialized by default\n\n## Provider execution model\n\nProvider calls are asynchronous work items. Vox must not treat correctness as \"only one provider request can exist at a time.\"\n\nTTS providers, especially remote API-backed providers, should support concurrent `synthesize` calls. A client may submit many independent utterances and await their results independently. Playback ordering is a caller concern, not a provider-execution constraint.\n\nASR has more physical-resource constraints because microphone capture may involve one input device, permissions, and ownership. That constraint belongs to capture/session coordination, not to the provider protocol itself. File transcription and provider inference can still be concurrent when the selected backend has capacity.\n\nCapacity should be explicit:\n\n- providers may advertise or be configured with max concurrency\n- Vox may apply per-provider or per-model backpressure when capacity is exhausted\n- backpressure should return a typed busy/capacity error or queue metadata, not silently impose a global mutex\n- telemetry should distinguish provider execution time from queue/wait time when queueing exists\n\n## Writing a provider\n\n### Rust ASR adapters\n\nRust is the default implementation language for new external ASR providers. Use the `vox-provider` crate in `providers/vox-provider/`. It owns JSON-RPC parsing, result envelopes, progress notifications, error responses, and stdout isolation. Implement the `AsrProvider` trait:\n\n| Adapter method | Return value |\n|---|---|\n| `models()` | Model info list; inspect local state without downloading or loading weights |\n| `install(model_id, progress)` | Installed model info; prepare assets without loading the model |\n| `preload(model_id, progress)` | Model info after loading weights into memory |\n| `transcribe(path, model_id)` | Text, model id, elapsed time, stage metrics, and optional word timings |\n\nUse the supplied progress callback during installation or warm-up. Return `ProviderError` for an expected failure. Start the process with `serve(adapter)`. The adapter remains alive to retain model state between calls.\n\nThe host permits one operation per process. Overlapping calls receive JSON-RPC error `-32001` (busy). Invalid input and adapter failures return errors without ending the process. On stdin EOF, the host finishes the accepted operation and exits. On Unix, native stdout output is redirected to stderr while a separate descriptor carries protocol responses.\n\n`providers/example/` is the runnable starter. It validates WAV input and returns sample text without an inference engine. `providers/whistle/` is the real engine adapter. Add a crate to `providers/Cargo.toml` for another engine. Use an existing audio conversion library and expose install and warm-up separately.\n\n```bash\ncargo run --manifest-path providers/Cargo.toml -p vox-provider-example\n# Write one JSON-RPC request per line, for example:\n# {\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"models\",\"params\":{}}\n```\n\nSwift continues to own the embedded Apple engine and the daemon. Other provider languages remain compatible with the same wire protocol.\n\n### Bundling a provider\n\nAdd a directory under `packages/cli/plugins/<plugin-id>/` with `bundle.json` and a compiled executable:\n\n```json\n{\n \"executables\": { \"darwin-arm64\": \"bin/darwin-arm64/vox-whistle\" },\n \"env\": { \"VOX_PROVIDER_CALL_TIMEOUT_SECONDS\": \"300\" }\n}\n```\n\nKeys identify the operating system and architecture used by Node (`darwin-arm64` for Apple Silicon). The installer selects the matching executable, copies the bundle into `VOX_HOME/plugins/<plugin-id>/`, and writes an absolute command path. Native executables must be regular files inside the bundle. Unsupported platforms and missing binaries fail before changing an existing installation. Plugin registration does not run the provider.\n\nBuild Whistle's Apple Silicon bundle on macOS:\n\n```bash\nrustup target add aarch64-apple-darwin\nbun run build:providers\nbun run --cwd packages/cli build\n```\n\nRust is required to build providers from source. Installed providers run without Rust, Python, or `uv`. The package release workflow builds the binary on macOS and includes it in the CLI package. The CLI's prepack check rejects packages that lack a declared provider executable.\n\nAdd a catalog plugin with `install: { \"kind\": \"bundle\", \"id\": \"<plugin-id>\" }` and a model entry whose `plugin` field matches the plugin id. Then run `vox plugins install <plugin-id>` and restart `voxd`. Extend `scripts/build-providers.ts` when adding a compiled bundle.\n\nExisting single-file `.mjs` bundles and interpreted directory bundles remain supported. Interpreted bundles use `command`, optional `shared` files from `plugins/shared/`, and `{pluginDir}` substitution. Their launcher is resolved using the installer's `PATH`.\n\nBundle manifests and files ship with the CLI. A remote catalog can select a shipped bundle; it cannot replace that bundle's command. Bundle paths, shared filenames, symlinks, and commands are validated before installation changes files.\n\nRun `bun run test:providers` to test the Rust protocol and adapters. These tests do not download model assets. Use a separate `VOX_HOME` for a real installation and transcription check.\n\n### Other runtimes\n\nA provider is any executable that reads newline-delimited JSON-RPC from stdin and writes responses to stdout. Minimal TypeScript example:\n\n```typescript\n// minimal-provider.ts\nimport { createInterface } from \"readline\";\n\nconst rl = createInterface({ input: process.stdin });\n\nfor await (const line of rl) {\n const req = JSON.parse(line);\n\n if (req.method === \"models\") {\n respond(req.id, {\n models: [\n {\n id: \"my-model:v1\",\n name: \"My Model\",\n backend: \"custom\",\n installed: true,\n preloaded: false,\n available: true,\n },\n ],\n });\n }\n\n if (req.method === \"transcribe\") {\n const text = await myTranscribe(req.params.path);\n respond(req.id, {\n modelId: req.params.modelId,\n text,\n elapsedMs: 100,\n metrics: { inferenceMs: 95, totalMs: 100 },\n });\n }\n\n if (req.method === \"voices\") {\n respond(req.id, {\n voices: [\n {\n id: \"default\",\n name: \"Default\",\n backend: \"custom-tts\",\n modelId: req.params?.modelId ?? \"my-tts:v1\",\n available: true,\n default: true,\n },\n ],\n });\n }\n\n if (req.method === \"synthesize\") {\n const audioBase64 = await mySynthesize(req.params.input);\n respond(req.id, {\n modelId: req.params.modelId,\n voiceId: req.params.voiceId ?? \"default\",\n format: \"wav\",\n contentType: \"audio/wav\",\n audioBase64,\n elapsedMs: 120,\n metrics: { synthesisMs: 110, totalMs: 120 },\n });\n }\n}\n\nfunction respond(id: number, result: unknown) {\n process.stdout.write(JSON.stringify({ jsonrpc: \"2.0\", id, result }) + \"\\n\");\n}\n```\n\nRegister it in `~/.vox/providers.json`:\n\n```json\n{\n \"providers\": [\n {\n \"id\": \"my-provider\",\n \"kind\": \"asr\",\n \"command\": [\"bun\", \"run\", \"minimal-provider.ts\"],\n \"models\": [\"my-model:v1\"]\n },\n {\n \"id\": \"my-tts\",\n \"kind\": \"tts\",\n \"command\": [\"bun\", \"run\", \"minimal-provider.ts\"],\n \"models\": [\"my-tts:v1\"]\n }\n ]\n}\n```\n\nThen select it via CLI or SDK by specifying the target model ID.\n\n## Provider lifecycle\n\nVox spawns the provider process on first use. It stays alive for the daemon's lifetime. If it crashes, Vox restarts it on the next request.\n\nProviders should be stateless between requests. The provider process can keep model weights in memory, but Vox assumes nothing about that state -- a crash and restart must not break anything."
},
{
"id": "sdk",
"title": "SDK (Companion Client)",
"description": "Typed Bun and Node access to Vox Companion over local WebSocket JSON-RPC.",
"content": "> Use `@voxd/sdk` when you want a Bun or Node tool to talk to `voxd` over local WebSocket JSON-RPC. For Apple apps, embed the Swift packages directly. For web apps or browser extensions, use [`@voxd/client`](./web-integration.md) instead.\n\n`packages/client/` connects to `voxd` when you want out-of-process access to models, voices, warm-up, transcription, synthesis, and stage metrics.\n\n`listCatalog()` and `refreshCatalog()` read the published dictation catalog. Plugin install stays on the CLI (`vox plugins install`). See [Models and plugins](./models.md).\n\n## Example\n\n```ts\nimport { VoxClient } from \"@voxd/sdk\";\n\nconst client = new VoxClient({ clientId: \"menu-bar\" });\n\nawait client.connect();\nawait client.scheduleWarmup(\"parakeet:v3\", 500);\n\nconst transcript = await client.transcribeFile(\"/tmp/sample.wav\", \"parakeet:v3\");\nconst voices = await client.listVoices(\"avspeech:system\");\nconst speech = await client.synthesize(\"Hello from Vox\", {\n modelId: \"avspeech:system\",\n voiceId: voices[0]?.id,\n format: \"wav\",\n});\n\nconsole.log(transcript.text);\nconsole.log(transcript.metrics?.inferenceMs);\nconsole.log(transcript.words);\nconsole.log(speech.audioBytes);\nconsole.log(speech.metrics?.synthesisMs);\n\nclient.disconnect();\n```\n\n## Client Identity\n\n`clientId` is used to attribute latency by consumer, compare route-level behavior across integrations, and support multi-client workflows.\n\n## Main methods\n\n```ts\ninterface VoxClientSurface {\n connect(): Promise<void>;\n disconnect(): void;\n doctor(): Promise<unknown>;\n listModels(): Promise<unknown>;\n listCatalog(): Promise<SpeechModelCatalog>;\n refreshCatalog(): Promise<SpeechModelCatalog>;\n listVoices(modelId?: string): Promise<unknown>;\n installModel(modelId?: string): Promise<unknown>;\n preloadModel(modelId?: string): Promise<unknown>;\n getWarmupStatus(modelId?: string): Promise<unknown>;\n startWarmup(modelId?: string): Promise<unknown>;\n scheduleWarmup(modelId?: string, delayMs?: number): Promise<unknown>;\n transcribeFile(path: string, modelId?: string): Promise<FileTranscriptionResult>;\n annotateFile(path: string, options?: AnnotationOptions): Promise<FileAnnotationResult>;\n synthesize(text: string, options?: SynthesisOptions): Promise<SynthesisResult>;\n getLiveSessionStatus(): Promise<LiveSessionStatus | null>;\n cancelLiveSession(sessionId?: string): Promise<{ cancelled: boolean; sessionId: string }>;\n createLiveSession(): VoxLiveSession;\n}\n```\n\n## File result shape\n\n```ts\ninterface FileTranscriptionResult {\n modelId: string;\n text: string;\n elapsedMs: number;\n metrics?: TranscriptionMetrics;\n words: WordTiming[];\n}\n```\n\n## Synthesis result shape\n\n```ts\ninterface SynthesisResult {\n modelId: string;\n voiceId: string;\n format: string;\n contentType: string;\n audio: Uint8Array;\n audioBytes: number;\n elapsedMs: number;\n metrics?: SynthesisMetrics;\n speechTiming?: SpeechTiming;\n}\n```\n\n## Error handling\n\nAll client methods throw when `voxd` is unreachable, the model isn't installed, or a transcription or synthesis request fails. Errors are plain `Error` instances, so check `message` for a human-readable description.\n\n```ts\ntry {\n const result = await client.transcribeFile(\"/tmp/audio.wav\");\n} catch (err) {\n // Common causes:\n // - Companion not running: start with `vox daemon start`\n // - Model not installed: run `vox models install` first\n // - Voice mismatch: inspect `client.listVoices(modelId)`\n // - Request failed: daemon logs have details (`vox logs daemon`)\n console.error(err.message);\n}\n```\n\nFor live sessions, call `session.cancel()` in a `finally` block to ensure the microphone is always released:\n\n```ts\nconst session = await client.createLiveSession();\ntry {\n // ...use session\n} finally {\n await session.cancel();\n}\n```\n\n## Configuration\n\n```ts\nconst client = new VoxClient({\n clientId: \"menu-bar\", // stable identity for telemetry\n port: 42137, // override the `companion-ws` daemon port\n host: \"127.0.0.1\", // override daemon host\n});\n```\n\nOn the daemon side, set `VOX_PORT` or `VOX_HOST` environment variables to override defaults. `VOX_PORT` controls the `companion-ws` daemon port discovered from `~/.vox/runtime.json`.\n\n## Integration advice\n\n- embed Swift directly for macOS and iOS apps; use `@voxd/sdk` when you want Vox Companion access from JS or tooling\n- use a stable `clientId` per product surface such as `menu-bar`, `browser-extension`, or `vox-cli`\n- warm on intent, not on every keystroke\n- call `listVoices(modelId)` before pinning a TTS voice in product code\n- benchmark with representative audio clips and read `inferenceMs` separately from `totalMs`\n- preserve raw transcription and synthesis metrics in your own telemetry if the app already exports traces"
},
{
"id": "web-integration",
"title": "Web Integration (Companion Client)",
"description": "Probe and use Vox Companion from web apps through the origin-gated local HTTP bridge.",
"content": "> For Apple apps, embed Vox's Swift packages directly. For Bun/Node companion clients, use [`@voxd/sdk`](./sdk.md) instead. It connects to `voxd` over local WebSocket JSON-RPC.\n\n`@voxd/client` lets a web app or browser extension talk to Vox Companion on the user's Mac over a local HTTP bridge. No server required.\n\nThis browser client is STT / alignment focused today. For TTS, use the companion-facing TypeScript SDK or the CLI.\n\n## Install\n\n```bash\nnpm install @voxd/client\n```\n\n## Quick start\n\n```ts\nimport { createVoxdClient } from \"@voxd/client\";\n\nconst client = createVoxdClient();\n\n// Check if the companion is running\nif (await client.probe()) {\n // Transcribe audio from a blob\n const result = await client.transcribe({\n audio: audioBlob,\n language: \"en\",\n timestamps: true,\n });\n\n console.log(result.text);\n console.log(result.words); // word-level timestamps\n}\n```\n\n## Discovery\n\nCall `probe()` on page load. It hits the companion's health endpoint with a short timeout and returns `true` or `false`. Fails silently when the companion is not installed.\n\n```ts\nconst client = createVoxdClient();\nconst available = await client.probe();\n```\n\nAfter probing, check `client.state` for the current connection state: `\"connected\"`, `\"unavailable\"`, `\"probing\"`, or `\"unknown\"`.\n\n## Capabilities\n\nOnce connected, check what the companion supports:\n\n```ts\nconst caps = await client.capabilities();\n\nif (caps.features.alignment) {\n // Word-level timestamps available\n}\n\nif (caps.features.local_asr) {\n // Local transcription available\n}\n```\n\n## Transcription\n\n### From a Blob or File\n\nUse `transcribe()` when you have audio data in the browser (recording, TTS clip, file upload).\n\n```ts\nconst result = await client.transcribe({\n audio: blob, // Blob, File, or ArrayBuffer\n language: \"en\",\n timestamps: true, // include word-level timing\n});\n\nresult.text; // full transcript\nresult.words; // [{ word, start, end }, ...]\nresult.durationMs; // audio duration\n```\n\n### From a URL\n\nUse `align()` when the audio lives on a server. The companion fetches it directly, avoiding a round trip through the browser.\n\n```ts\nconst alignment = await client.align({\n source: {\n audioUrl: \"https://your-app.com/api/audio/abc123\",\n format: \"mp3\",\n },\n metadata: {\n documentId: \"doc_123\",\n pageNumber: 2,\n },\n});\n\nalignment.words; // [{ word, start, end }, ...]\nalignment.durationMs;\n```\n\n`align()` creates a job, polls until done, and returns the result. Blocks up to 5 minutes.\n\n### Lower-level job API\n\nFor more control, use `createJob()` and `getJob()` directly:\n\n```ts\nconst { jobId } = await client.createJob({\n type: \"alignment\",\n source: { audioUrl: \"https://your-app.com/audio/abc.mp3\" },\n metadata: { cacheKey: \"abc123\" },\n});\n\n// Poll manually\nconst status = await client.getJob(jobId);\n// status.status: \"accepted\" | \"processing\" | \"completed\" | \"failed\"\n// status.result?.alignment: { words, durationMs }\n```\n\n## Fallbacks\n\nVox Companion will not be installed or running on every machine. It helps to probe for it and keep a fallback path ready when it is unavailable.\n\n```ts\nconst client = createVoxdClient();\n\nasync function getAlignment(audioUrl: string) {\n // Try local companion first\n if (await client.probe()) {\n try {\n return await client.align({ source: { audioUrl } });\n } catch {\n // Fall through to cloud\n }\n }\n\n // Fallback to cloud API or heuristic timing\n return await cloudAlignmentFallback(audioUrl);\n}\n```\n\n## When the companion isn't installed\n\nIf `probe()` returns false, you can prompt the user to install:\n\n```ts\nif (!await client.probe()) {\n // Show install prompt in your UI\n // Link to: https://voxd.cc/download\n}\n```\n\nOr try launching it via deep link (works if installed but not running):\n\n```ts\nclient.launch(); // triggers vox://launch\n```\n\n## Error handling\n\nAll methods throw `VoxDError` with a `code` property:\n\n```ts\nimport { VoxDError } from \"@voxd/client\";\n\ntry {\n const result = await client.transcribe({ audio: blob });\n} catch (err) {\n if (err instanceof VoxDError) {\n switch (err.code) {\n case \"network_error\": // companion unreachable\n case \"http_error\": // non-2xx response\n case \"job_failed\": // transcription failed\n case \"timeout\": // job took too long\n case \"no_result\": // job completed without result\n }\n }\n}\n```\n\n## HTTP bridge reference\n\nThe companion HTTP bridge listens on `http://127.0.0.1:43115` by default (the `companion-http` port, configurable via `host` and `port` options). These endpoints are what `@voxd/client` calls under the hood.\n\n| Method | Path | Auth | Description |\n|--------|------|------|-------------|\n| `GET` | `/health` | Open | Liveness check |\n| `GET` | `/capabilities` | Origin | Features, backends, models |\n| `POST` | `/jobs` | Origin | Create alignment/transcription job |\n| `GET` | `/jobs/:id` | Origin | Poll job status |\n| `POST` | `/transcribe` | Origin | Upload audio for transcription |\n| `GET` | `/live` | Origin | Live session status |\n| `POST` | `/live` | Origin | Start a live recording session (streaming NDJSON) |\n| `POST` | `/live/stop` | Origin | Stop a live session and get final transcript |\n| `POST` | `/live/cancel` | Origin | Cancel a live session without transcribing |\n\n**Origin gating:** All endpoints except `/health` require a valid `Origin` header. Vox ships with built-in origins for first-party apps. Add your own in Vox settings, or drop a JSON file into `~/.vox/origins.d/`:\n\n```json\n{\"origins\":[\"https://app.example.com\"]}\n```\n\nVox merges all origin sources. Wildcard ports work on loopback hosts (`http://localhost:*`).\n\n## Configuration\n\n```ts\nconst client = createVoxdClient({\n host: \"127.0.0.1\", // default; override for non-loopback setups\n port: 43115, // override the `companion-http` bridge port\n baseUrl: \"http://...\",// overrides host + port when set\n clientId: \"my-app\", // stable identity for telemetry\n probeTimeout: 2000, // ms before probe gives up\n pollInterval: 500, // ms between job status polls\n});\n```\n\nOn the daemon side, set `VOX_PORT`, `VOX_BRIDGE_PORT`, or `VOX_HOST` environment variables to override defaults. `VOX_BRIDGE_PORT` controls the `companion-http` bridge, while `VOX_PORT` controls the underlying `companion-ws` daemon."
},
{
"id": "observability",
"title": "Observability",
"description": "Read Vox performance samples, stage timings, client dimensions, and operator dashboards.",
"content": "Telemetry is built into the runtime for both transcription and synthesis. Each performance sample includes:\n\n- `clientId`\n- `route`\n- `modelId`\n- `voiceId` for synthesis routes when a voice is selected\n- `outcome`\n- nested `metrics`\n\n## Metrics\n\nTranscription metrics:\n\n- `fileCheckMs`\n- `modelCheckMs`\n- `modelLoadMs`\n- `audioLoadMs`\n- `audioPrepareMs`\n- `inferenceMs`\n- `totalMs`\n- `audioDurationMs`\n\nSynthesis metrics:\n\n- `modelCheckMs`\n- `modelLoadMs`\n- `voiceResolveMs`\n- `synthesisMs`\n- `totalMs`\n- `audioDurationMs`\n- `outputBytes`\n- `characterCount`\n\nDerived values: `realtimeFactor`, warm vs cold from `modelLoadMs`, audio-to-text speed from `audioDurationMs / inferenceMs`, and text-to-audio speed from `audioDurationMs / synthesisMs`.\n\n## Storage\n\nThe runtime appends JSON lines to:\n\n```text\n~/.vox/performance.jsonl\n```\n\nThe CLI dashboard reads from this file. You can also export it to another metrics backend.\n\nCurrent emitted performance routes:\n\n- `transcribe.file`\n- `transcribe.live`\n- `synthesize.generate`\n- `synthesize.startSession`\n\nCancelled synthesis sessions are recorded under `synthesize.startSession` with `outcome: \"cancelled\"`.\n\n## Operator Commands\n\n```bash\nvox transcribe file --metrics /tmp/sample.wav\nvox transcribe bench /tmp/sample.wav 5\nvox speak --metrics \"hello world\"\nvox speak bench \"hello world\" 5\nvox perf dashboard\nvox perf dashboard --client vox-cli\n```\n\n## Reading the numbers\n\n`inferenceMs`, `synthesisMs`, and `totalMs` measure different things.\n\n- For ASR, `inferenceMs` is how fast the hot model ran.\n- For TTS, `synthesisMs` is how long audio generation took once the request was inside the model.\n- `totalMs` is what the user experienced end-to-end.\n\n## Example sample\n\n```json\n{\n \"timestamp\": \"2026-04-25T17:54:26Z\",\n \"clientId\": \"menu-bar\",\n \"route\": \"synthesize.generate\",\n \"modelId\": \"avspeech:system\",\n \"voiceId\": \"com.apple.voice.compact.en-US.Samantha\",\n \"outcome\": \"ok\",\n \"textLength\": 14,\n \"metrics\": {\n \"audioDurationMs\": 1240,\n \"synthesisMs\": 182,\n \"inferenceMs\": 182,\n \"totalMs\": 196\n }\n}\n```\n\n## Dashboard tips\n\n- Only compare clients when the prompt or audio shape is similar.\n- Use `inferenceMs` for loaded-model ASR speed.\n- Use `synthesisMs` for loaded-model TTS speed.\n- Use `totalMs` for end-user latency.\n- Large `modelLoadMs` spikes are warm-up events, not inference regressions.\n\nSee [Runtime](./runtime.md) for emitted routes and [API](./api.md) for the current performance sample shape."
},
{
"id": "architecture",
"title": "Architecture",
"description": "Layer ownership and data flow across Swift engines, companion transports, TypeScript clients, and operator tools.",
"content": "## Layers\n\n### VoxCore\n\nShared runtime types and utilities:\n\n- runtime metadata\n- transcription and synthesis metrics\n- performance samples\n- filesystem paths\n- model catalog types\n- plugin install records\n- trace utilities\n\n### VoxEngine\n\nModel-facing speech layer:\n\n- model installation and preload\n- ASR provider routing and audio preparation\n- annotation provider routing and speaker-attribution contracts\n- TTS provider routing and voice discovery\n- Parakeet inference\n- OpenAI file transcription\n- catalog ingest for known ASR families\n- AVSpeech, OpenAI, ElevenLabs, MiniMax, NVIDIA Magpie, Groq, Gemini, and external synthesis backends\n- stage-level timing\n\n### VoxAppleSpeech\n\nOptional Apple embed playback layer:\n\n- per-audible-surface speech output controller\n- live system delivery for `avspeech:system`\n- generated-audio playback through an injectable sink\n- typed playback phases without app product policy\n\n### VoxService\n\nDaemon-side orchestration:\n\n- JSON-RPC bridge\n- annotation route dispatch\n- live session coordination\n- synthesis session coordination\n- microphone recording\n- warm-up scheduling\n- performance sample recording\n\n### TypeScript SDK\n\n`@voxd/sdk`: health, models, voices, warm-up, file transcription, synthesis, live sessions, metrics parsing.\n\n### Browser SDK\n\n`@voxd/client`: probe, transcribe, align, live sessions over the HTTP bridge.\n\n### Companion bridge\n\n`VoxBridge` / `voxbridge`: browser-facing HTTP bridge that proxies into the companion daemon while keeping the browser surface narrower than the underlying WebSocket RPC runtime.\n\n### CLI\n\n`@voxd/cli`: operator tool. Doctor, daemon lifecycle, model catalog, plugins, voices, transcription, synthesis, benchmarks, dashboards.\n\n## Ownership\n\n| Surface | Owns |\n|---------|------|\n| Swift runtime | Daemon lifecycle, audio prep, model lifecycle, provider routing, transcription, annotation, synthesis, perf recording |\n| VoxAppleSpeech | Optional per-audible-surface Apple playback: live system speech, generated-audio sinks, replace/stop/cancel, typed phase events |\n| TypeScript SDK | Connection lifecycle, typed request/response shapes, live-session ergonomics, transcription and synthesis metric parsing |\n| Browser SDK | Companion discovery, audio upload, job polling, live sessions over HTTP bridge |\n| CLI | Operator commands, terminal output (human and machine), warm-up controls, transcription, synthesis, dashboards |\n| Site and docs | Architecture docs, onboarding, OG images, landing page |\n\n## Data flow\n\n1. Client creates a connection with a stable `clientId`\n2. CLI or SDK issues JSON-RPC to `voxd`, while browser clients reach the same runtime through `VoxBridge`\n3. `VoxService` coordinates model state and route dispatch\n4. `VoxEngine` prepares ASR input, annotation input, or TTS requests and dispatches them to the selected provider\n5. `VoxCore` types and trace utilities shape the result\n6. Runtime appends tagged performance samples for local inspection\n\n```text\nApple app ───────────────▶ VoxEngine\nBun/Node tool ─▶ voxd ──▶ VoxService ─▶ VoxEngine\nBrowser ───────▶ VoxBridge ─▶ voxd ───▶ VoxService ─▶ VoxEngine\n```\n\nSee [Runtime](./runtime.md) for route-level behavior and [Provider Protocol](./providers.md) for external engine boundaries."
},
{
"id": "api",
"title": "API",
"description": "Current companion RPC routes, TypeScript SDK entrypoints, result types, and performance shapes.",
"content": "Public protocol and SDK-facing type shapes.\n\n## RPC Methods\n\n### Health and Runtime\n\n- `health`\n- `doctor.run`\n\n### Models\n\n- `models.list`\n- `models.install`\n- `models.preload`\n- `models.catalog`\n- `models.refreshCatalog`\n\nCatalog JSON may include `plugins[]`. Plugin install is a CLI file-system action (`vox plugins install`), not an RPC. `voxd` loads `~/.vox/plugins/*/provider.json` at start.\n\n### Warm-Up\n\n- `warmup.status`\n- `warmup.start`\n- `warmup.schedule`\n\n### Transcription\n\n- `transcribe.file`\n- `transcribe.startSession`\n- `transcribe.sessionStatus`\n- `transcribe.stopSession`\n- `transcribe.cancelSession`\n\n### Annotation\n\n- `annotate.file`\n\n### History\n\n- `history.list`\n- `history.get`\n- `history.delete`\n- `history.clear`\n\n### Synthesis\n\n- `synthesize.voices`\n- `synthesize.generate`\n- `synthesize.startSession`\n- `synthesize.sessionStatus`\n- `synthesize.cancel`\n\n## Performance samples\n\nThese fields are present on every performance sample recorded to `~/.vox/performance.jsonl`.\n\n```ts\ntype PerformanceRoute =\n | \"annotate.file\"\n | \"transcribe.file\"\n | \"transcribe.live\"\n | \"synthesize.generate\"\n | \"synthesize.startSession\"\n | string;\n```\n\n```ts\ninterface PerformanceSample {\n timestamp: string;\n clientId: string;\n route: PerformanceRoute;\n modelId: string;\n voiceId?: string;\n outcome: \"ok\" | \"error\" | \"cancelled\" | string;\n textLength: number;\n error?: string;\n metrics?: PerformanceMetrics;\n}\n\ninterface PerformanceMetrics {\n traceId: string;\n audioDurationMs: number;\n wasPreloaded: boolean;\n modelCheckMs: number;\n modelLoadMs: number;\n inferenceMs: number;\n totalMs: number;\n inputBytes?: number;\n fileCheckMs?: number;\n audioLoadMs?: number;\n audioPrepareMs?: number;\n characterCount?: number;\n outputBytes?: number;\n voiceResolveMs?: number;\n synthesisMs?: number;\n realtimeFactor?: number;\n}\n```\n\nCurrent emitted performance routes:\n\n- `transcribe.file`\n- `annotate.file`\n- `transcribe.live`\n- `synthesize.generate`\n- `synthesize.startSession`\n\n## Core TypeScript SDK Entry Points\n\n### `VoxClient`\n\n- `connect()`\n- `disconnect()`\n- `doctor()`\n- `listModels()`\n- `listCatalog()`\n- `refreshCatalog()`\n- `listVoices()`\n- `installModel()`\n- `preloadModel()`\n- `getWarmupStatus()`\n- `startWarmup()`\n- `scheduleWarmup()`\n- `transcribeFile()`\n- `annotateFile()`\n- `synthesize()`\n- `getLiveSessionStatus()`\n- `cancelLiveSession()`\n- `createLiveSession()`\n\n### `FileTranscriptionResult`\n\n- `modelId`\n- `text`\n- `elapsedMs`\n- `metrics`\n- `words`\n\n### `SynthesisResult`\n\n- `modelId`\n- `voiceId`\n- `format`\n- `contentType`\n- `audio`\n- `audioBytes`\n- `elapsedMs`\n- `metrics`\n\n### `FileAnnotationResult`\n\n- `modelId`\n- `text`\n- `elapsedMs`\n- `metrics`\n- `words`\n- `speakers`\n\n### `TranscriptionMetrics`\n\n- `traceId`\n- `audioDurationMs`\n- `inputBytes`\n- `wasPreloaded`\n- `fileCheckMs`\n- `modelCheckMs`\n- `modelLoadMs`\n- `audioLoadMs`\n- `audioPrepareMs`\n- `inferenceMs`\n- `totalMs`\n- `realtimeFactor`\n\n### `SynthesisMetrics`\n\n- `traceId`\n- `characterCount`\n- `audioDurationMs`\n- `outputBytes`\n- `wasPreloaded`\n- `modelCheckMs`\n- `modelLoadMs`\n- `voiceResolveMs`\n- `synthesisMs`\n- `inferenceMs`\n- `totalMs`\n- `realtimeFactor`\n\n### `AnnotationMetrics`\n\n- `traceId`\n- `audioDurationMs`\n- `inputBytes`\n- `wasPreloaded`\n- `fileCheckMs`\n- `modelCheckMs`\n- `modelLoadMs`\n- `audioLoadMs`\n- `audioPrepareMs`\n- `diarizationMs`\n- `totalMs`\n- `realtimeFactor`\n\n## Interface shapes\n\n```ts\ninterface TranscriptionMetrics {\n traceId: string;\n audioDurationMs?: number;\n inputBytes?: number;\n wasPreloaded?: boolean;\n fileCheckMs?: number;\n modelCheckMs?: number;\n modelLoadMs?: number;\n audioLoadMs?: number;\n audioPrepareMs?: number;\n inferenceMs?: number;\n totalMs?: number;\n realtimeFactor?: number;\n}\n\ninterface FileTranscriptionResult {\n modelId: string;\n text: string;\n elapsedMs: number;\n metrics?: TranscriptionMetrics;\n words: WordTiming[];\n}\n\ninterface SpeakerSegment {\n speakerId: string;\n start: number;\n end: number;\n confidence?: number | null;\n}\n\ninterface AttributedWordTiming {\n word: string;\n start: number;\n end: number;\n confidence: number;\n speakerId?: string | null;\n}\n\ninterface AnnotationMetrics {\n traceId: string;\n audioDurationMs: number;\n inputBytes: number;\n wasPreloaded: boolean;\n fileCheckMs: number;\n modelCheckMs: number;\n modelLoadMs: number;\n audioLoadMs: number;\n audioPrepareMs: number;\n diarizationMs: number;\n totalMs: number;\n realtimeFactor: number;\n}\n\ninterface FileAnnotationResult {\n modelId: string;\n text?: string;\n elapsedMs: number;\n metrics?: AnnotationMetrics;\n words: AttributedWordTiming[];\n speakers: SpeakerSegment[];\n}\n\ninterface SynthesisOptions {\n modelId?: string;\n voiceId?: string;\n format?: string;\n speed?: number;\n instructions?: string;\n speechTiming?: boolean | SpeechTimingRequest;\n credentials?: {\n OPENAI_API_KEY?: string;\n openaiApiKey?: string;\n openai_api_key?: string;\n NV_API_KEY?: string;\n NVIDIA_API_KEY?: string;\n nvApiKey?: string;\n nvidiaApiKey?: string;\n nv_api_key?: string;\n nvidia_api_key?: string;\n GROQ_API_KEY?: string;\n groqApiKey?: string;\n groq_api_key?: string;\n GEMINI_API_KEY?: string;\n GOOGLE_API_KEY?: string;\n GOOGLE_GENAI_API_KEY?: string;\n geminiApiKey?: string;\n googleApiKey?: string;\n googleGenaiApiKey?: string;\n gemini_api_key?: string;\n google_api_key?: string;\n google_genai_api_key?: string;\n };\n}\n\ninterface SpeechTimingRequest {\n enabled?: boolean;\n modelId?: string;\n cues?: Array<{\n id: string;\n text?: string;\n textStart?: number;\n textEnd?: number;\n }>;\n}\n\ninterface SynthesisMetrics {\n traceId: string;\n characterCount: number;\n audioDurationMs: number;\n outputBytes: number;\n wasPreloaded: boolean;\n modelCheckMs: number;\n modelLoadMs: number;\n voiceResolveMs: number;\n synthesisMs: number;\n inferenceMs: number;\n totalMs: number;\n realtimeFactor: number;\n}\n\ninterface SynthesisResult {\n modelId: string;\n voiceId: string;\n format: string;\n contentType: string;\n audio: Uint8Array;\n audioBytes: number;\n elapsedMs: number;\n metrics?: SynthesisMetrics;\n speechTiming?: SpeechTiming;\n}\n\ninterface SpeechTiming {\n source: \"asr\" | \"native\" | \"estimated\" | string;\n modelId: string;\n elapsedMs: number;\n words: SpeechTimingWord[];\n cues: SpeechTimingCue[];\n}\n```\n\n## Warm-up states\n\n```ts\ntype WarmupState = \"idle\" | \"scheduled\" | \"warming\" | \"ready\" | \"failed\";\n```\n\nApps use this to tell whether the runtime is cold, warming, or ready for hot-path speech.\n\nSee the [SDK Guide](./sdk.md) for usage patterns and [Runtime](./runtime.md) for route ownership and session semantics."
},
{
"id": "agents",
"title": "Agent Guide",
"description": "Canonical routing, source-of-truth, and verification guidance for agents working with Vox.",
"content": "Use this page to choose the right Vox surface before editing code or proposing an integration.\n\n## Choose the integration mode first\n\n| Caller | Use | Do not use by default |\n|---|---|---|\n| macOS or iOS app process | `VoxCore` + `VoxEngine` (+ optional `VoxAppleSpeech`) | `voxd`, `@voxd/sdk`, `@voxd/client` |\n| Bun or Node tool | `voxd` + `@voxd/sdk` | browser bridge |\n| Web app or browser extension | Vox Companion + `@voxd/client` | direct WebSocket RPC |\n| Operator or benchmark workflow | `@voxd/cli` | custom scripts that hide metrics |\n\nWarm-up is always explicit. Preserve `clientId`, `route`, `modelId`, and `voiceId` where applicable.\n\n## Canonical source map\n\n| Question | Read first |\n|---|---|\n| Swift package products and platforms | `swift/Package.swift` |\n| ASR/TTS provider behavior | `swift/Sources/VoxEngine/` |\n| Optional Apple speech output | `swift/Sources/VoxAppleSpeech/` |\n| Full Vox companion app | `apps/vox/` |\n| Minivox dictation app | `apps/minivox/` |\n| Runtime routes and session ownership | `swift/Sources/VoxService/VoxRuntimeService.swift` |\n| Companion SDK methods and types | `packages/client/src/client.ts`, `packages/client/src/types.ts` |\n| Browser API and HTTP paths | `packages/web-client/src/client.ts`, `packages/web-client/src/types.ts` |\n| CLI syntax | `packages/cli/src/index.ts` and `vox help` |\n| Dictation catalog and plugins | `data/models.json` and [Models and plugins](./models.md) |\n| Public integration guidance | the matching page in `docs/` |\n\nWhen prose and source disagree, treat the source as current and update the prose in the same change.\n\n## Agent briefs\n\nMajor docs pages have a paired compact brief under `docs/agent/*.agent.md`. In the rendered docs, **Copy agent brief** copies that paired file. **Copy markdown** always copies the full human guide.\n\nUse `llms.txt` for routing and `llms-full.txt` when a complete documentation handoff is needed.\n\n## Verification by surface\n\n```bash\nbun test packages/client/test packages/cli/test packages/web-client/test\n\nswift test --package-path swift\n\nbun run site:build\nbun run docs:check\n```\n\nPrefer the narrowest relevant command while iterating. Run the full surface check before publishing.\n\n## Do not assume\n\n- Apple apps need Vox Companion.\n- Embed mode records telemetry automatically.\n- Warm-up happens implicitly.\n- Stop and cancel mean the same thing.\n- A copied type shape is current without checking package exports and source types.\n- A proposal document describes shipped behavior unless its implementation is confirmed in source.\n\nStart with the [Overview](./overview.md), then open the integration guide selected by the table above."
},
{
"id": "skill",
"title": "Operations Guide",
"description": "Repeatable health, warm-up, benchmark, and performance-triage workflows for operators and agents.",
"content": "Highest-value workflows:\n\n- add warm-up before first speech\n- preserve `clientId` in SDK initialization\n- benchmark warm performance with real audio\n- read the local dashboard before speculating about latency\n\nRecommended operator loop:\n\n```bash\nvox doctor\nvox models catalog\nvox warmup start\nvox transcribe bench /path/to/audio.wav 5\nvox perf dashboard --client <integration>\n```\n\n1. `vox doctor`\n2. `vox models catalog`\n3. `vox warmup start`\n4. `vox transcribe bench /path/to/audio.wav 5`\n5. `vox perf dashboard --client <integration>`\n\nTo add Gemma 4 or another catalog plugin: `vox plugins install mlx-vlm`, then restart `voxd`.\n\n## Recommended client naming\n\nUse stable product-surface IDs instead of per-user or per-session IDs:\n\n- `vox-cli`\n- `menu-bar`\n- `browser-extension`\n- `editor-plugin`\n\nThis keeps dashboard slices meaningful over time.\n\n## Performance triage order\n\nWhen a user reports that transcription feels slow:\n\n1. confirm whether the report is about hot-path inference or cold-path readiness\n2. inspect `inferenceMs` before speculating about model quality\n3. inspect `totalMs` and `modelLoadMs` to separate warm-up cost from steady-state cost\n4. compare samples by `clientId` and `route`\n\n## Contributor checklist\n\n- keep `clientId`, `route`, and `modelId` intact in telemetry\n- avoid hiding model lifecycle inside helpers that make latency opaque\n- prefer repeatable file-based benchmarks before changing live-session behavior\n\nSee [Observability](./observability.md) for metric interpretation and [Quickstart](./quickstart.md) for installation recovery."
}
],
"agentBriefs": [
{
"id": "agents",
"title": "Agent Routing Facts",
"content": "# Agent Routing Facts\n\n- choose integration mode before choosing packages\n- Apple app process: `VoxCore` + `VoxEngine` (+ optional `VoxAppleSpeech`)\n- Bun/Node tool: `voxd` + `@voxd/sdk`\n- browser: Vox Companion + `@voxd/client`\n- operator workflow: `@voxd/cli`\n- canonical SDK types: `packages/client/src/types.ts`\n- canonical browser surface: `packages/web-client/src/client.ts`\n- canonical CLI syntax: `packages/cli/src/index.ts`\n- canonical runtime dispatch: `swift/Sources/VoxService/VoxRuntimeService.swift`\n- update docs in the same change when source and prose disagree\n- preserve explicit warm-up and telemetry dimensions"
},
{
"id": "api",
"title": "API Facts",
"content": "# API Facts\n\n- do not copy public type shapes without checking `packages/client/src/types.ts`\n- companion RPC dispatch is canonical in `swift/Sources/VoxService/VoxRuntimeService.swift`\n- current file routes: `transcribe.file`, `annotate.file`, `synthesize.generate`\n- warm-up routes: `warmup.status`, `warmup.start`, `warmup.schedule`\n- live transcription has start, status, stop, and cancel routes\n- synthesis has voices, generate, session start, status, and cancel routes\n- `createLiveSession()` is synchronous and returns `VoxLiveSession`\n- synthesis options include optional speech timing and allowlisted OpenAI/NVIDIA/Groq/Gemini credentials; ElevenLabs and MiniMax keys are not lent per-request\n- performance samples preserve `clientId`, `route`, `modelId`, and optional `voiceId`"
},
{
"id": "apple-embed",
"title": "Apple Embed Facts",
"content": "# Apple Embed Facts\n\n## Use embed mode when\n\n- target: macOS app or iOS app\n- caller: app process\n- requirement: voice in and voice out inside the app\n\n## Use companion mode when\n\n- target: web app or browser extension\n- caller: Bun/Node tool outside the app process\n- requirement: shared local voice service across processes\n\n## Do not do this\n\n- do not start `voxd` inside the Apple app just to call Vox APIs\n- do not use `@voxd/sdk` or `@voxd/client` from app code\n- do not hide warm-up as an implicit side effect\n\n## Add package dependency\n\n- remote dependency: `.package(url: \"https://github.com/hudsonkit/vox.git\", from: \"0.5.2\")`\n- sibling checkout dependency: `.package(path: \"../vox/swift\")`\n- required products for app embed: `VoxCore`, `VoxEngine`\n- optional product for Apple playback: `VoxAppleSpeech` / `AppleSpeechOutputController`\n- avoid `VoxService` and `VoxBridge` unless intentionally embedding companion/runtime behavior\n\n## Start with VoxDictation\n\n- `VoxDictation(clientId:modelId:preferredInputDeviceID:engine:recordsPerformance:)`, an actor in `VoxEngine`\n- `warmUp(progress:)` downloads and loads the model; call on user intent, not at launch\n- `start(onBuffer:)` records the microphone; macOS only, throws on iOS\n- `stop()` returns `TranscriptionOutput` and deletes the recording; `cancel()` discards\n- `transcribe(fileURL:)` works on macOS and iOS\n- `inputLevel()` for meters; `isListening`, `isReady`\n- errors: `VoxDictationError.alreadyListening`, `.notListening`\n- records a `PerformanceSample` per transcription with routes `transcribe.dictation` / `transcribe.file`\n\n## Default embed engines\n\n- ASR default: `EngineManager()` -> `ParakeetProvider()`\n- TTS generation default: `TTSEngineManager()` -> `TTSProviderRegistry` with every built-in provider (OpenAI, ElevenLabs, MiniMax, NVIDIA, Groq, Gemini, AVSpeech); routes by model id, never falls back between providers\n- TTS playback is not on `TTSProvider`; optional Apple playback is `AppleSpeechOutputController`\n- default ASR model id: `parakeet:v3`\n- English-only ASR model id: `parakeet:v2`\n- default TTS model id: `TTSDefaults.modelId` = `gpt-4o-mini-tts`\n- local TTS model id: `TTSDefaults.localModelId` = `avspeech:system`\n- default TTS format: `TTSDefaults.format` = `wav`\n\n## Public types to use\n\n- ASR input: `URL`\n- ASR output: `TranscriptionOutput`\n- TTS request: `SynthesisRequest`\n- TTS output: `SynthesisOutput`\n- voices: `TTSVoiceInfo`\n- telemetry: `PerformanceRecorder`, `PerformanceSample`\n\n## Warm-up\n\n- dictation: call `dictation.warmUp(progress:)`\n- call `asr.preload(modelId:progress:)`\n- call `tts.preload(modelId:voiceId:progress:)`\n- do not rely on `WarmupCoordinator`; it is not public in the embed surface\n\n## Telemetry parity\n\n- `VoxDictation` records automatically; `EngineManager` and `TTSEngineManager` callers must record manually\n- preserve fields: `clientId`, `route`, `modelId`, `voiceId`\n- preserve route names: `transcribe.dictation`, `transcribe.file`, `synthesize.generate`\n- performance log path comes from `RuntimePaths.performanceLogURL()`\n\n## App owns these concerns\n\n- microphone permission (`NSMicrophoneUsageDescription`; audio-input entitlement when sandboxed)\n- audio capture on iOS (macOS can use `VoxDictation` / `MicrophoneFileRecorder`)\n- product-level spoken-output policy (dedupe, markdown flattening, fallback copy, preferences, queue priority)\n- playback of `SynthesisOutput.audioData`, or an opt-in `AppleSpeechOutputController` per audible surface\n- interruption handling\n- product state and UI\n\n## OpenAI TTS rule\n\n- `gpt-4o-mini-tts` needs an OpenAI key; without one the request throws, with no automatic fallback\n- use `TTSDefaults.localModelId` for the on-device system voice; retry with it yourself if you want a fallback\n- key lookup order: `SynthesisRequest.providerCredentials`, `ProviderEntry.env`, process env, Vox credential store\n- pass `OPENAI_API_KEY` via `ProviderEntry.env` or `providerCredentials`; never ship a long-lived key in the binary\n- do not rely on process environment inside iOS app code\n\n## Optional Apple playback\n\n- `VoxAppleSpeech` is optional and per audible surface, not a singleton\n- `avspeech:system` uses live `AVSpeechSynthesizer.speak()`, never `write()`-then-play\n- generated-audio models play `SynthesisOutput.audioData` through an injectable player sink\n- audio-session configuration is injectable and off by default\n- new requests replace pending generation/playback\n- stop/cancel are idempotent and must cancel Task, player, and synthesizer during the enqueue window\n- events report resolving/generating, starting, playing, finished, cancelled, failed\n- synthesis identity reports requested vs actual model/voice, not a guessed provider id\n- synthesis identity is separate from the physical audio-output route\n- do not put reply dedupe, markdown flattening, fallback copy, preference storage, queue priority, or telemetry double-recording in this controller\n\n## Known gaps\n\n- `VoxDictation` covers dictation only; `VoxAppleSpeech` is separate playback\n- no microphone capture on iOS yet\n- no public embed live-session coordinator yet (no partial text while recording)\n- no raw-buffer ASR API yet\n- no automatic telemetry recording outside `VoxDictation`\n- `VoxAppleSpeech` does not own browser playback or app product policy\n\n## Default recipe\n\n- client id: stable per app, for example `my-app-ios` / `my-app-macos`\n- dictation: `VoxDictation`\n- ASR engine for custom flows: `EngineManager()`\n- TTS engine: `TTSEngineManager()`\n- TTS default: `gpt-4o-mini-tts`\n- TTS local model: `avspeech:system` (fallback is the app's choice)\n- use Vox Companion only for web or cross-process workflows\n\n## Packaged app resources\n\nXcode copies SwiftPM resource bundles automatically. When packaging by hand, copy the bundle SwiftPM built into the signed app's `Contents/Resources`: `Vox_HudsonSpeechEngine.bundle` when depending on the repo root, `HudsonSpeechEngine_HudsonSpeechEngine.bundle` when depending on `vox/swift`. Vox looks for both.\n`SpeechEngineResources` resolves the model catalog and mlx-audio provider script\nfrom that location in apps, and from SwiftPM resources in command-line builds.\nA packaged app does not fall back to a developer build directory."
},
{
"id": "architecture",
"title": "Architecture Facts",
"content": "# Architecture Facts\n\n- `VoxCore`: shared types, paths, traces, and performance records\n- `VoxEngine`: provider routing, model lifecycle, audio preparation, ASR, annotation, and TTS (AVSpeech, OpenAI, ElevenLabs, MiniMax, NVIDIA Magpie, Groq, Gemini, external)\n- `VoxAppleSpeech`: optional per-audible-surface Apple playback controller (live system speech, generated-audio sinks, replace/stop/cancel, typed phase events)\n- `VoxService`: daemon RPC, sessions, warm-up, capture, and telemetry\n- `VoxBridge`: narrow browser-facing HTTP transport\n- `@voxd/sdk`: WebSocket companion client for Bun/Node\n- `@voxd/client`: HTTP browser client\n- `@voxd/cli`: operator and benchmark surface\n- Swift owns the embeddable engine and companion transport surfaces\n- keep model lifecycle and warm-up explicit across layers"
},
{
"id": "models",
"title": "Models and plugin facts",
"content": "# Models and plugin facts\n\n- catalog URL: `https://voxd.cc/data/models.json`\n- repo source: `data/models.json`\n- bundled fallback: `swift/Sources/HudsonSpeechEngine/Resources/models.json`\n- cache: `~/.vox/cache/models-catalog.json`\n- override: `VOX_MODEL_CATALOG_URL`\n- refresh: `vox models catalog` / `vox models catalog refresh`; RPC `models.catalog` / `models.refreshCatalog`\n- SDK: `listCatalog()`, `refreshCatalog()`\n- catalog refresh never installs or runs a plugin\n- default ASR: `parakeet:v3`; English TDT: `parakeet:v2`\n- Apple Silicon only; Intel Macs are unsupported\n- native core ASR: `apple:speech-transcriber` on macOS 26+\n- Apple locale env: `VOX_APPLE_SPEECH_LOCALE`\n- recommended MLX candidates: `mlx-community/Qwen3-ASR-1.7B-8bit`, `mlx-community/cohere-transcribe-03-2026-mlx-8bit`, `mlx-community/nemotron-3.5-asr-streaming-0.6b-8bit`\n- catalog capability flags describe the current Vox provider surface; `liveTranscription=false` until the ASR protocol and mic path emit model-backed partials\n- never add hand-written resampling or denoising inside a model adapter; use platform/provider DSP libraries\n- OpenAI file ASR: `gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `whisper-1`\n- plugins CLI: `vox plugins list|install|remove`\n- plugin store: `~/.vox/plugins/<id>/provider.json`\n- `voxd` loads installed plugins at start\n- Gemma 4 E2B: model `gemma-4-e2b-it`, plugin `mlx-vlm`\n- plugin launchers: `node`, `bun`, `npx`, `bunx`, `uv`, `uvx`, `python3`, `python`; install resolves launchers to absolute paths\n- Whistle: model/plugin `whistle`, family `cactus-whistle`, optional Rust provider using prebuilt Needle CPU runtime; no Needle text model\n- Whistle setup: bundled Apple Silicon executable; `vox plugins install whistle`, restart voxd, `vox models install whistle`, `vox models preload whistle`\n- provider runs without Python/uv/Rust installed; only model install downloads weights/runtime; model listing does not load weights\n- Whistle: file ASR up to 30 seconds, word timestamps/confidence, one operation per process; overlapping calls return busy\n- Whistle adapter uses Symphonia/Rubato for audio conversion; `liveTranscription=false`\n- Whistle assets: `VOX_HOME/models/whistle`; transcribe can cold-load installed assets but cannot download them\n- Whistle env: `VOX_WHISTLE_LANGUAGE` (en/de/fr/es/it/nl/pl, unset auto), `VOX_WHISTLE_KEYWORDS` (comma-separated)"
},
{
"id": "observability",
"title": "Observability Facts",
"content": "# Observability Facts\n\n- primary operator commands:\n - `vox transcribe file --metrics /path/to/audio.wav`\n - `vox transcribe bench /path/to/audio.wav 5`\n - `vox speak --no-play --output /tmp/out.wav \"hello world\"`\n - `vox speak bench \"hello world\" 5`\n - `vox perf dashboard`\n- inference speed should be evaluated from `inferenceMs` / `synthesisMs`, not `totalMs`\n- total latency should be evaluated from `totalMs`\n- exported dimensions:\n - `clientId`\n - `route`\n - `modelId`\n - `voiceId` for synthesis"
},
{
"id": "overview",
"title": "Vox Facts",
"content": "# Vox Facts\n\n- platforms: `macOS 14+` / `iOS 17+` for Swift transcription embedding; `macOS 26+` for Minivox and the Hudson menu app\n- deployment modes:\n - embed mode: Swift packages inside the app process\n - companion mode: `voxd` for web, browser, and shared-process clients\n- default ASR embed engine: `EngineManager()` -> `ParakeetProvider()`\n- default TTS embed engine: `TTSEngineManager()` -> `TTSProviderRegistry`\n- optional Apple TTS playback: `VoxAppleSpeech` / `AppleSpeechOutputController`\n- companion service: `VoxRuntimeService`\n- CLI: `@voxd/cli` (Bun)\n- companion SDK: `@voxd/sdk` (TypeScript, WebSocket JSON-RPC to `voxd`)\n- browser SDK: `@voxd/client` (HTTP bridge to Vox Companion)\n- full companion app: `apps/vox/`\n- Minivox dictation app: `apps/minivox/`\n- Minivox install: `npx -y @voxd/cli@latest install mini` or `brew install --cask arach/vox/minivox`\n- Minivox first use: the installer launches the menu-bar app; put the cursor in a text field, press `Right ⌘M` to start, then press `Right ⌘M` again to stop and copy/paste the transcript\n- Minivox command: `minivox`, `minivox settings`, `minivox quit`\n- Minivox npm install output: `--quiet` or `--verbose`\n- model focus: `parakeet:v3` for ASR, `parakeet:v2` for English-only TDT, `gpt-4o-mini-tts` for default TTS\n- dictation catalog: `https://voxd.cc/data/models.json` and site `/models`; refresh with `vox models catalog refresh`\n- plugins: catalog `plugins[]`; install with `vox plugins install <id>`; stored in `~/.vox/plugins/`\n- warmup: explicit API surface\n- telemetry dimensions: `clientId`, `route`, `modelId`, `voiceId`\n- configurable: `VOX_PORT`, `VOX_BRIDGE_PORT`, `VOX_HOST`, `VOX_HOME`, `VOX_MODEL_CATALOG_URL`"
},
{
"id": "providers",
"title": "Provider Protocol Facts",
"content": "# Provider Protocol Facts\n\n- transport: newline-delimited JSON-RPC 2.0 over stdin/stdout\n- register ASR and TTS as separate `providers.json` entries\n- `command` is a string array containing executable and arguments\n- common methods: `models`, `install`, `preload`\n- install/preload results wrap model info as `{ \"model\": { ... } }`\n- ASR method: `transcribe` with an audio file `path`\n- TTS methods: `voices`, `synthesize`\n- ASR results require `metrics.inferenceMs` and `metrics.totalMs`\n- TTS results require `metrics.totalMs`; prefer `synthesisMs`\n- builtin TTS ids: `avspeech`, `openai-tts`, `elevenlabs`, `minimax`, `nvidia`, `groq`, `gemini`\n- NVIDIA Magpie: model `magpie-tts-multilingual`, `NV_API_KEY`/`NVIDIA_API_KEY`, LINEAR_PCM 44.1 kHz, dynamic `list_voices`, reject >2000 chars\n- Groq Orpheus: WAV `response_format`, 200-character limit, `GROQ_API_KEY`, require structurally valid WAV\n- Gemini TTS: Generate Content AUDIO for 2.5 and `gemini-3.1-flash-tts-preview`, wrap `audio/L16` as WAV, `GEMINI_API_KEY`/`GOOGLE_API_KEY`, voice `Rasalgethi`\n- per-request `credentials` allowlist: OpenAI, NVIDIA, Groq, Gemini/Google only; never ElevenLabs/MiniMax; never log keys\n- explicit empty/whitespace provider env key or alias fences process credential fallback for resolution and default model selection; process env is used only when that key is absent from `env`\n- NVIDIA Magpie, Groq, and Gemini synthesis HTTP errors redact the exact request text; JSON diagnostics keep secret redaction; plain-text bodies fail closed whenever synthesis sensitive values are supplied; voice discovery has no prompt; older providers keep their existing error surfaces\n- remote providers participate in defaults only when configured; aliases such as `magpie`/`groq-tts`/`google-tts` block canonical default overwrite\n- default TTS ranking is independent of config and registry order: OpenAI `gpt-4o-mini-tts`, ElevenLabs, MiniMax, NVIDIA, Groq, Gemini, then AVSpeech\n- reserve stdout for protocol messages and send logs to stderr\n- provider state may cache weights, but crash/restart must remain safe\n- canonical implementations: `swift/Sources/HudsonSpeechEngine/` built-in providers plus `ExternalProvider.swift` and `ExternalTTSProvider.swift`\n- published catalog: `https://voxd.cc/data/models.json` from repo `data/models.json`\n- catalog families the engine can run: `parakeet-tdt`, `apple-speech`, `mlx-audio`, `openai-transcribe`\n- catalog plugins: `plugins[]` plus model `plugin` id; install writes `~/.vox/plugins/<id>/provider.json`\n- Gemma 4 E2B: plugin `mlx-vlm`; `vox plugins install mlx-vlm`\n- Whistle: optional `whistle` plugin/model, prebuilt Needle runtime in a separate Rust process; file ASR <=30 seconds, word timings, no live partials\n- default external ASR implementation: Rust; shared `vox-provider` crate in `providers/vox-provider/`, implement `AsrProvider`, call `serve(adapter)`\n- Rust host: one active operation, `-32001` busy, progress callback, Unix stdout descriptor isolation; EOF drains accepted work\n- starter: `providers/example/`; real adapter: `providers/whistle/`; workspace: `providers/Cargo.toml`\n- Rust tests: `bun run test:providers`; macOS build: `rustup target add aarch64-apple-darwin`, `bun run build:providers`\n- installed providers need no Rust/Python/uv; Swift still owns embedded Apple engines and daemon\n- native directory bundles: `plugins/<id>/bundle.json` executables map (`darwin-arm64` to bundle-relative binary), optional args/env\n- interpreted directory bundles: command/env/shared; `{pluginDir}` substitution; shared files from `plugins/shared/`\n- installer validates before writes, confines native executable to bundle, preserves executable mode, resolves interpreted launchers to absolute paths; legacy single `.mjs` supported\n- release builds native artifacts on macOS; CLI prepack rejects missing binaries\n- refresh never executes plugin commands\n- plugin launchers allowlist: `node`, `bun`, `npx`, `bunx`, `uv`, `uvx`, `python3`, `python`\n- refresh: `models.catalog` / `models.refreshCatalog`; CLI `vox models catalog [refresh]`\n- plugins CLI: `vox plugins list|install|remove`\n- default local ASR: `parakeet:v3`; English TDT: `parakeet:v2`\n- native alternative: `apple:speech-transcriber`\n- MLX shortlist: Qwen3-ASR 1.7B, Cohere Transcribe 03-2026, Nemotron 3.5 ASR Streaming 0.6B\n- supported architecture: Apple Silicon (`arm64`) only; no Intel support\n- current provider `transcribe` is file-based; do not claim model-backed partials from upstream streaming capability\n- model adapters use platform/provider audio conversion; never add hand-written resampling or denoising\n- remote OpenAI ASR: `gpt-transcribe`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, `whisper-1`\n- override catalog URL with `VOX_MODEL_CATALOG_URL`"
},
{
"id": "quickstart",
"title": "Quickstart Facts",
"content": "# Quickstart Facts\n\n- Swift embed minimums: macOS 14+, iOS 17+\n- Minivox direct-embed dictation app: `apps/minivox/`, macOS 26+\n- packaged Vox menu app requires macOS 26+\n- published CLI requires Node 22+\n- install CLI: `npm install -g @voxd/cli`\n- install Companion from the DMG before `vox install`, or provide `~/.vox/bin/voxd`\n- verify with `vox doctor`\n- repo checkout CLI after build: `node packages/cli/dist/index.js`\n- ASR smoke path: `vox warmup start parakeet:v3` then `vox transcribe file <path>`\n- catalog: `vox models catalog` then `vox models catalog refresh`\n- plugins: `vox plugins list` then `vox plugins install mlx-vlm`\n- do not send Apple-native integrations through Companion by default"
},
{
"id": "runtime",
"title": "Runtime Facts",
"content": "# Runtime Facts\n\n- health methods: `health`, `doctor.run`\n- warmup methods: `warmup.status`, `warmup.start`, `warmup.schedule`\n- ASR model routes: `models.list`, `models.install`, `models.preload`, `models.catalog`, `models.refreshCatalog`\n- ASR file route: `transcribe.file`\n- annotation route: `annotate.file`\n- ASR live routes: `transcribe.startSession`, `transcribe.sessionStatus`, `transcribe.stopSession`, `transcribe.cancelSession`\n- TTS routes: `synthesize.voices`, `synthesize.generate`, `synthesize.startSession`, `synthesize.sessionStatus`, `synthesize.cancel`\n- performance log path: `~/.vox/performance.jsonl`\n- runtime metadata path: `~/.vox/runtime.json`\n- model catalog cache: `~/.vox/cache/models-catalog.json`\n- plugin store: `~/.vox/plugins/<id>/provider.json`\n- browser bridge endpoints other than `/health` are origin-gated\n- session ownership includes both `connectionID` and `clientId`; stop and cancel are not interchangeable"
},
{
"id": "sdk",
"title": "SDK Facts",
"content": "# SDK Facts\n\n- companion client entrypoint: `packages/client/src/client.ts`\n- metrics parser: `packages/client/src/metrics.ts`\n- companion client only: Apple apps should embed Swift packages directly instead\n- core result types:\n - `FileTranscriptionResult`\n - `FileAnnotationResult`\n - `SynthesisResult`\n- voice metadata type: `VoiceInfo`\n- warmup methods exposed: `getWarmupStatus`, `startWarmup`, `scheduleWarmup`\n- synthesis methods exposed: `listVoices`, `synthesize`\n- ASR method exposed: `transcribeFile`\n- catalog methods exposed: `listCatalog`, `refreshCatalog`\n- annotation method exposed: `annotateFile`\n- live session method exposed: `createLiveSession`; it returns `VoxLiveSession` synchronously\n- synthesis can return optional `speechTiming`; read the types in `packages/client/src/types.ts` before duplicating shapes\n- per-request `credentials` are allowlisted for OpenAI, NVIDIA Magpie, Groq, and Gemini/Google only"
},
{
"id": "skill",
"title": "Operations Facts",
"content": "# Operations Facts\n\n- start with `vox doctor`\n- distinguish cold readiness from hot inference before changing code\n- warm explicitly with `vox warmup start [modelId]`\n- benchmark ASR with `vox transcribe bench <path> [runs]`\n- benchmark TTS with `vox speak bench <text> [runs]`\n- inspect tagged samples with `vox perf dashboard`\n- keep client IDs stable by product surface, not user or session\n- read `inferenceMs` or `synthesisMs` separately from `totalMs`\n- use file-based benchmarks before changing live-session behavior"
},
{
"id": "web-integration",
"title": "Web Integration Facts",
"content": "# Web Integration Facts\n\n- package: `@voxd/client`\n- transport: local HTTP bridge, default `http://127.0.0.1:43115`\n- call `probe()` before depending on Companion availability\n- use `transcribe()` for browser-owned audio data\n- use `align()` when Companion should fetch an audio URL\n- live sessions use `createLiveSession()` and distinct stop/cancel semantics\n- browser client does not own microphone capture or permissions\n- all bridge endpoints except `/health` are origin-gated\n- preserve a stable `clientId`\n- use `launch()` only as an installed-but-not-running recovery path\n- canonical surface: `packages/web-client/src/client.ts`"
}
]
}