ActiveVision
PulseAugur coverage of ActiveVision — every cluster mentioning ActiveVision across labs, papers, and developer communities, ranked by signal.
-
GPT-5.5 and Claude Fable 5 fail new ActiveVision benchmark, scoring far below humans
A new benchmark called ActiveVision, designed to test repeated visual perception, has revealed significant limitations in advanced AI models. GPT-5.5 achieved only a 10.6% success rate, failing entirely on 11 out of 17 …
-
New ActiveVision benchmark reveals MLLMs lack human-like visual observation
A new benchmark called ActiveVision has been developed to test the active observation capabilities of multimodal large language models (MLLMs). This benchmark, comprising 17 tasks, aims to measure how well these models …
-
New ActiveVision benchmark reveals MLLMs lack human-like active observation
A new benchmark called ActiveVision has been developed to test the active observation capabilities of multimodal large language models (MLLMs), a crucial aspect of human vision that involves continuous redirection of ga…