Researchers have introduced RoleBreak, a new benchmark designed to evaluate the long-horizon role-playing capabilities of spoken dialogue systems. The benchmark includes over 300 roles and thousands of human-verified dialogue turns, with specific criteria for assessing role consistency, interaction quality, safety, and vocal emotion over extended conversations. Evaluations of nine system configurations revealed that while current models are better at maintaining semantic roles than vocal emotion, they struggle with long-term consistency, failing on persona and safety after an average of around 11 turns. Scaling the language model significantly improved semantic robustness but had little effect on vocal expressiveness, indicating persistent gaps in spoken role-playing systems. AI
IMPACT Highlights persistent challenges in maintaining persona consistency and vocal expressiveness in long-duration spoken dialogue systems.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- RoleBreak
- ScienceCast
- scite Smart Citations
- Text To Speech
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →