Two new benchmarks, M3-DuplexBench and Hy-MultiTurn, have been released to evaluate the capabilities of multi-turn dialogue models. M3-DuplexBench focuses on full-duplex spoken dialogue systems, supporting English and Japanese across various domains and dialogue contexts. Hy-MultiTurn, designed for Chinese dialogue understanding, assesses models on six specific dimensions including constraint memory and reference resolution over extended conversations. AI
IMPACT These benchmarks will drive improvements in multi-turn dialogue systems, pushing models to better handle complex conversations and diverse languages.
RANK_REASON Two new academic papers introducing novel benchmarks for evaluating dialogue models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →