Researchers have developed a new method for synthesizing non-verbal vocalizations (NVs) in text-to-speech (TTS) systems, focusing on preference optimization techniques. They introduced an NV-aware character error rate (NV-CER) to control the realization of NVs like laughter and coughs without altering the core optimization algorithm. Experiments on the Emilia-NV dataset and the augmented NV-Bench demonstrated the effectiveness of their approach, providing practical guidance for improving expressive TTS. AI
IMPACT Enhances expressiveness in TTS systems, potentially leading to more natural and engaging synthetic voices.
RANK_REASON The cluster contains a research paper detailing a new method for synthesizing non-verbal vocalizations in TTS systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Direct Preference Optimization
- Emilia-NV
- Gotit.pub
- Hugging Face
- NV-Bench
- NV-CER
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →