Researchers have developed TontaubeV1, a novel text-to-speech model designed to balance natural prosody with efficient streaming capabilities. The model utilizes a hierarchical DualCodec representation, separating semantic and acoustic information to predict utterance duration and refine audio. TontaubeV1 can process up to one minute of reference audio for voice conditioning and achieves a streaming latency of approximately 200ms on a single RTX 5090 GPU. Benchmarks indicate its prosody quality matches ElevenLabs Flash v2.5 and surpasses other leading models. AI
IMPACT This model's streaming capabilities and prosody quality could advance real-time voice applications and content generation.
RANK_REASON The item describes a new model release with a research paper and released weights. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- DualCodec
- ElevenLabs Flash v2.5
- Fish Audio S2 Pro
- Gradium API
- Hugging Face
- Qwen3 0.6B
- Qwen3 1.7B
- RTX 5090
- TontaubeV1
- VibeVoice
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →