A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
-
Updated
Sep 6, 2026 - Python
A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
An Open-Source Project to Unify Audio Processing and Generation
A 1.6B causal Transformer audio tokenizer with streaming, variable bitrates, and semantic alignment across speech, sound, and music
[Official Implementation] Acoustic Autoregressive Modeling 🔥
An objective reconstruction evaluation toolkit for audio tokenizers, neural codecs, audio VAEs, and vocoders
Unofficial PyTorch implementation of Higgs Audio V2 Tokenizer with HuBERT semantic features. Complete training pipeline for semantic-acoustic audio tokenization with 960x downsampling and 8-layer RVQ.
Neural audio codec and tokenizer for audio language models — SEANet encoder with residual vector quantization in PyTorch
低比特率神经音频编解码器 + 残差矢量量化:把连续音频转成离散 token,服务音频语言模型
Audio-token compression benchmark for speech ML infra: EnCodec + ASR/Speaker evals, Modal GPU runs, salience/trained selectors, KV-cache serving analysis, and public benchmark artifacts.
Run private RAG, knowledge bases, and AI coding tools on your hardware using Ollama, FastGPT, and pgvector. Keep your data local with no cloud dependencies.
To associate your repository with the audio-tokenizer topic, visit your repo's landing page and select "manage topics."