Researchers have introduced new datasets and models focused on full-duplex spoken dialogue systems, which enable more natural, real-time conversational interactions. The DuplexDrama dataset offers over 2,000 hours of synthesized audio data covering scenarios, expressive speech, and sound events, with a subset of bilingual dialogues to be released. SteerDuplex introduces a model and benchmark for steerable full-duplex speech, allowing control over attributes like tone and speaking rate, and demonstrating significant improvements in instruction following and interruption handling. Additionally, ConversationalVoice presents a pipeline to create training data from real conversations, generating separated, reconstructed, and expanded speech artifacts that preserve interaction dynamics. AI
IMPACT These advancements in full-duplex spoken dialogue systems could lead to more natural and responsive AI assistants and conversational agents.
RANK_REASON Multiple research papers introducing new datasets and models for spoken dialogue systems.
- ConversationalVoice
- Gemini
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DuplexDrama
- Gotit.pub
- Hugging Face
- Moshi
- ScienceCast
- SteerBench
- SteerDuplex
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →