A new solver for the ARC-AGI-2 visual reasoning benchmark has achieved a top score of 72.9% on the semi-private evaluation set, outperforming leading frontier models like GPT-5.2 Pro and Gemini 3 Pro. The solver employs a modality-driven search strategy, generating reasoning candidates across text, image, and code, and uses a holistic judging approach to compare these candidates within a single prompt. This method effectively identifies correct minority hypotheses, even when the primary modal answer is incorrect. The research also highlights that prescriptive prompting and iterative refinement can reduce hypothesis diversity and degrade performance. AI
IMPACT Sets a new benchmark for visual reasoning, potentially influencing future LLM evaluation and development strategies.
RANK_REASON Research paper detailing a new solver for a benchmark with performance claims against existing models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →