Researchers have introduced DataVista, the first benchmark designed to evaluate multimodal large language models (MLLMs) on their ability to understand data videos. This benchmark includes 961 real-world data videos and over 6,700 questions, structured to assess capabilities in data perception, temporal reasoning, and narrative understanding. Evaluations of 19 MLLMs revealed that even the top performer, Gemini-3.1 Pro, achieved only 70.0% accuracy, significantly below human expert levels, with notable weaknesses in causal reasoning and narrative comprehension. The study also found that increasing frame counts and adding subtitles offered limited improvements, particularly for narrative understanding, and identified common failure modes in chart reading and evidence judgment. AI
IMPACT Highlights limitations in current MLLMs for specialized data interpretation tasks, indicating a need for improved multimodal reasoning capabilities.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- DataVista
- Gemini-3.1 Pro
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →