A new benchmark called ActiveVision has been developed to test the active observation capabilities of multimodal large language models (MLLMs). This benchmark, comprising 17 tasks, aims to measure how well these models can redirect their gaze based on intermediate hypotheses, similar to human vision. Early evaluations show that even top-tier models like GPT-5.5 and Claude Fable-5 perform poorly, solving only a small percentage of tasks, significantly underperforming human participants who averaged over 96%. The research suggests that current MLLMs lack robust active visual observation, highlighting a need for architectural and training objective improvements to better close the perception-reasoning loop. AI
IMPACT Highlights a critical gap in current MLLMs, potentially guiding future research towards more human-like visual perception and reasoning.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- ActiveVision
- alphaXiv
- arXiv
- CatalyzeX
- Claude Fable-5
- DagsHub
- Gotit.pub
- GPT-5.5
- Hugging Face
- human
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →