Researchers have developed TimeBlind, a new benchmark designed to test the spatio-temporal understanding capabilities of video Large Language Models (LLMs). The benchmark uses a minimal-pairs paradigm, presenting videos that are visually identical but differ only in their temporal structure, to isolate temporal reasoning from static visual cues. Evaluations show that even advanced models like GPT-5 and Gemini 3 Pro perform poorly, achieving only 48.2% accuracy compared to human performance of 98.2%, indicating a significant reliance on visual shortcuts rather than true temporal logic. AI
IMPACT Highlights a critical gap in current video LLM capabilities, likely driving future research towards more robust temporal reasoning.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →