Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
-
Updated
Aug 8, 2026
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
A New End-to-end Framework for Evaluating Voice Agents
🇺🇦 Open Source Ukrainian Text-to-Speech datasets
A Docker-based OpenAI-compatible Text-to-Speech API server powered by Kyutai's TTS models with GPU acceleration support.
Just a simple multimodal avatar interaction platform
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.
A unified benchmarking framework for evaluating Voice AI agents across conversational quality, audio realism, latency metrics, and safety guardrails with scalable multi-language stress testing.
Open-source real-time Voice AI infrastructure in Go. Stream audio via WebRTC or WebSocket, connect STT → LLM → TTS pipelines, and build scalable voice agents and conversational AI applications.
🇺🇦 Ukrainian RAD-TTS++ models (decoder + models with 3 voices) and HiFiGAN model
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
A source-linked directory of free and trial LLM APIs, multimodal models, embeddings, speech, translation, safety, and other inference endpoints. Companion catalog for freellmapi.io.
Legacy Speech AI examples with migration links to the current Brainiall TTS and transcription services.
A curated list of the best Text-to-Speech, speech synthesis, and voice-cloning research — models, papers, benchmarks, and toolkits, focused on 2025–2026.
中文 ASR 评测工具箱 · micro-CER 对比 FunASR/Whisper/llama.cpp · 一条命令出报告 · 自带迷你测试集 · Mandarin ASR benchmark toolkit
Code-switching ASR adaptation for strong multilingual speech recognition models. Synthetic CSW data generation, Whisper adaptation, Bayesian LoRA (BLoRA), and robust multilingual ASR evaluation.
Interruptible voice-agent runtime for structured interview prototypes, with VAD-based interruption handling and modular speech backends.
Lucida Harness is a lightweight test harness and local Web UI for evaluating background removal models and alpha matting algorithms.
Add a description, image, and links to the speech-ai topic page so that developers can more easily learn about it.
To associate your repository with the speech-ai topic, visit your repo's landing page and select "manage topics."