A new benchmark, DuplexSpeechBench-IFEval (DSB-IFEval), has been introduced to evaluate how well full-duplex voice agents can implicitly follow instructions based on roles or personas, rather than explicit commands. The benchmark includes 1,038 test cases across eight assistant roles and measures instruction adherence and persona consistency. Initial testing revealed that while models like F-Actor and PersonaPlex are sensitive to instruction type, others like GPT-Realtime and MiniCPM-o excel at persona consistency but struggle to adapt their floor behavior across different instruction methods. AI
IMPACT This benchmark could lead to more natural and intuitive voice agent interactions by improving their ability to infer user intent from context.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- DSB-IFEval
- DuplexSpeechBench-IFEval
- F-Actor
- Fun-Audio-Chat
- gpt-realtime
- Instruction Adherence Score
- MiniCPM-o
- Persona Adherence Score
- PersonaPlex
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →