Researchers have introduced Instruct-FD, a new benchmark designed to evaluate the ability of full-duplex speech systems to follow turn-taking instructions. This is crucial for real-world applications where conversational policies need to adapt. The benchmark utilizes a synthetic pipeline for generating instruction-conditioned conversations and an LLM-based judge for evaluation. Initial testing on six state-of-the-art systems revealed a significant gap in instruction-following capabilities, with the best model achieving only 64.4% adherence, particularly struggling with proactive behaviors like backchanneling and interruption. AI
IMPACT Highlights a critical gap in current full-duplex speech systems, indicating a need for improved adaptability and instruction-following capabilities for real-world deployment.
RANK_REASON The item is a research paper introducing a new benchmark and evaluation protocol for AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Instruct-FD
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →