Researchers have introduced ReTurn, a new benchmark designed to evaluate how well multimodal AI models selectively use conversational history. The benchmark consists of 7,000 tasks that test the models' ability to retain relevant historical questions while discarding outdated answers, or to use historical evidence for current questions despite conflicting new information. Initial evaluations across 13 models showed a significant drop in accuracy when models transitioned from direct input to conversational settings, indicating challenges in effectively managing historical context. AI
IMPACT This benchmark could drive improvements in how AI models handle context and memory in conversational settings.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →