A new benchmark called StateSight has been introduced to evaluate the spatial-state reconstruction capabilities of vision-language models. The benchmark includes three tasks: cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Human participants significantly outperformed both OpenAI GPT-5.5 and Claude Sonnet 5 on these tasks, highlighting the models' difficulties in accurately reconstructing spatial information. The results suggest that models can produce format-valid responses that mask underlying failures in visual inference. AI
IMPACT Highlights limitations in current vision-language models for spatial reasoning, potentially guiding future research and development.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →