A new benchmark called MotionBlind has been developed to test the motion understanding capabilities of Video Large Language Models (Video-LLMs). Researchers found that most open-source Video-LLMs perform poorly, often failing to distinguish between different speeds or directions of motion, even when presented with clear visual data. While Gemini-3.1 Pro showed some improvement, it still struggled with accurately assessing speed, indicating that current Video-LLMs are not yet reliable for tasks requiring a true understanding of motion. AI
IMPACT Highlights critical limitations in current Video-LLMs, suggesting they are not yet suitable for applications requiring nuanced motion perception.
RANK_REASON Research paper introducing a new benchmark for evaluating Video-LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →