A new benchmark called PolyComp has been introduced to test the compositional 3D spatial reasoning capabilities of multimodal AI models. The benchmark consists of 120 procedurally generated problems, each requiring a model to identify two polycube components that form a target solid. In evaluations, GPT-5.6 Sol achieved 50.0% accuracy, while Claude Fable-5 reached 39.4%, and Gemini 3.1 Pro Preview scored 27.5%, which is close to the random guessing baseline. AI
IMPACT This benchmark could drive improvements in AI's ability to understand and reason about 3D spatial relationships, crucial for robotics and augmented reality applications.
RANK_REASON The cluster describes a new academic benchmark for AI model evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →