A new benchmark called JigShape has been developed to evaluate the visual-geometric reasoning capabilities of Vision-Language Models (VLMs). The benchmark uses interlocking puzzle pieces to create unambiguous ground truth, unlike previous methods with rectangular cuts. Initial testing revealed that most frontier models, including GPT-5.5, struggle significantly with zero-shot geometric reasoning, performing at chance levels on even small puzzles. While supervised fine-tuning improves performance on simpler grids, all models collapse on larger puzzles, indicating a current architectural limitation in maintaining constraint satisfaction as complexity increases. AI
IMPACT Highlights a significant gap in current VLMs' ability to perform complex geometric reasoning, suggesting a need for architectural advancements.
RANK_REASON Academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →