Researchers have developed a new benchmark, Autoregressive Mosaics (AM-Bench), to evaluate the 2D spatial reasoning capabilities of text-only large language models. The benchmark includes a translation task where models generate code for fully specified geometries and a layout task that assesses their ability to compose images from underspecified prompts. Results indicate that while all tested models can translate specified geometry into code, their performance on open-ended layout tasks varies significantly, suggesting differences beyond mere code-generation ability. The study also found that using raw SVG as an output medium improves layout scores compared to procedural code, and that models develop a coarse layout plan before generation, which evolves during the process. AI
IMPACT This research introduces a new method for evaluating LLM spatial reasoning, potentially guiding future model development towards better visual understanding.
RANK_REASON The cluster contains an academic paper detailing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- AM Bench 2022 Macroscale Tensile Challenge at Different Orientations (CHAL-AMB2022-04-MaTTO)
- arXiv
- Autoregressive Mosaics
- SVG
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →