Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
-
Updated
Jul 31, 2026 - Python
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
本地优先的 AI 视频字幕工作台:WhisperX/MLX 精准时间轴、混合语言与多人识别、对话感知双语翻译。
Point of Interest Error Rate (PIER) Metric for Code-Switching ASR: A specialized evaluation metric designed to focus on critical points in multilingual speech recognition, providing a more accurate analysis of code-switched utterances.
Real-time transformer-based ASR supporting 100+ languages - Google Cloud integration with noise cancellation & low-latency optimization
AISRT - 本地 AI 字幕生成工具 / local AI subtitle generator for video/audio to SRT, multilingual ASR, timestamp alignment, GUI/CLI batch processing, and local SRT translation.
Add a description, image, and links to the multilingual-asr topic page so that developers can more easily learn about it.
To associate your repository with the multilingual-asr topic, visit your repo's landing page and select "manage topics."