PulseAugur
EN
LIVE 01:10:04

New ActiveVision benchmark reveals MLLMs lack human-like visual observation

A new benchmark called ActiveVision has been developed to test the active observation capabilities of multimodal large language models (MLLMs). This benchmark, comprising 17 tasks, aims to measure how well these models can redirect their gaze based on intermediate hypotheses, similar to human vision. Early evaluations show that even top-tier models like GPT-5.5 and Claude Fable-5 perform poorly, solving only a small percentage of tasks, significantly underperforming human participants who averaged over 96%. The research suggests that current MLLMs lack robust active visual observation, highlighting a need for architectural and training objective improvements to better close the perception-reasoning loop. AI

IMPACT Highlights a critical gap in current MLLMs, potentially guiding future research towards more human-like visual perception and reasoning.

RANK_REASON The item is a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New ActiveVision benchmark reveals MLLMs lack human-like visual observation

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger ·

    An Exam for Active Observers

    arXiv:2607.16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wi…