Researchers have developed LoopTTS, a novel closed-loop Text-to-Speech (TTS) system designed to correct prosodic defects in synthesized audio. This system utilizes an AudioLLM as a judge to identify issues like misplaced stress or unnatural pauses, then employs a fine-grained instruction-following TTS model, the Refiner, to re-synthesize the audio with guided corrections. A new dataset, Refiner-DB, containing 42,000 annotated examples, was created to train the Refiner, demonstrating improved audio quality and better control over prosody compared to existing methods. AI
IMPACT This closed-loop TTS system could lead to more natural and expressive synthesized speech, improving accessibility and user experience in various applications.
RANK_REASON The cluster contains a research paper detailing a new system and dataset for TTS. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- AudioLLM
- CatalyzeX
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- LoopTTS
- Refiner-DB
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →