A new benchmark called ActiveVision, designed to test repeated visual perception, has revealed significant limitations in advanced AI models. GPT-5.5 achieved only a 10.6% success rate, failing entirely on 11 out of 17 tasks. Claude Fable 5 performed even worse with a 3.5% score, starkly contrasting with human participants who averaged 96.1% accuracy. The study highlights that even high-tier models struggle with tasks requiring continuous visual processing and cannot compensate by generating their own code. AI
IMPACT Highlights current AI limitations in continuous visual perception tasks, indicating a need for new architectures beyond current reasoning and coding capabilities.
RANK_REASON New benchmark evaluation of existing frontier models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →