A new benchmark called VideoZeroBench has been introduced to evaluate the spatio-temporal evidence verification capabilities of video multimodal large language models (V-MLLMs). This benchmark features manually annotated question-answer pairs across 13 video domains, focusing on fine-grained cues and distributed evidence. The evaluation protocol includes a five-level diagnostic system that assesses not only answer correctness but also the accuracy of temporal and spatial localization of evidence. Results show that while models like Gemini-3.7-Flash achieve moderate accuracy on standard question answering, their performance plummets when precise evidence localization is required, highlighting significant challenges in current V-MLLM systems. AI
IMPACT Highlights critical limitations in current video LLMs regarding evidence localization, driving future research towards more precise spatio-temporal understanding.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →