Researchers have introduced UltraVoice, a large-scale dataset designed to improve fine-grained speech style control in spoken dialogue models. The dataset includes over 830 hours of speech dialogues with instructions across six stylistic dimensions: emotion, speed, volume, accent, language, and composite styles. Fine-tuning models like SLAM-Omni and VocalNet on UltraVoice has shown significant improvements in stylistic controllability and instruction following, without compromising core conversational abilities. The dataset's utility also extends to training controllable Text-to-Speech models. AI
IMPACT Enhances human-like interaction in spoken dialogue systems and improves controllable Text-to-Speech models.
RANK_REASON The cluster contains an academic paper detailing a new dataset and methodology for speech synthesis. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →