A new research paper published on arXiv highlights significant limitations in current video language models, specifically Gemini 3.6 Flash. The study found that these models struggle with accurately counting events in videos, especially when events are transient or occur at high frequencies. While Gemini 3.6 Flash can reliably count persistent state transitions up to 12 events at a moderate frequency, it fails to accurately count blinking events and shows poor performance in high-count, high-frequency scenarios. The research suggests that the way events are represented and the temporal reasoning capabilities of these models are key bottlenecks, even when increasing video sampling rates or altering prompting strategies. AI
IMPACT Highlights critical limitations in video language models' temporal reasoning and event bookkeeping capabilities, suggesting current models are not ready for complex real-world video analysis tasks.
RANK_REASON Research paper detailing model limitations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →