PulseAugur
EN
LIVE 12:02:21

New solver beats GPT-5.2 Pro and Gemini 3 Pro on ARC-AGI-2 benchmark

A new solver for the ARC-AGI-2 visual reasoning benchmark has achieved a top score of 72.9% on the semi-private evaluation set, outperforming leading frontier models like GPT-5.2 Pro and Gemini 3 Pro. The solver employs a modality-driven search strategy, generating reasoning candidates across text, image, and code, and uses a holistic judging approach to compare these candidates within a single prompt. This method effectively identifies correct minority hypotheses, even when the primary modal answer is incorrect. The research also highlights that prescriptive prompting and iterative refinement can reduce hypothesis diversity and degrade performance. AI

IMPACT Sets a new benchmark for visual reasoning, potentially influencing future LLM evaluation and development strategies.

RANK_REASON Research paper detailing a new solver for a benchmark with performance claims against existing models.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New solver beats GPT-5.2 Pro and Gemini 3 Pro on ARC-AGI-2 benchmark

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Johan Land ·

    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

    arXiv:2606.31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I prese…

  2. arXiv cs.AI TIER_1 English(EN) · Johan Land ·

    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

    Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a solver for ARC-AGI-2, a few-shot visual rea…