PulseAugur
EN
LIVE 22:46:03

New solver beats GPT-5.2 Pro and Gemini 3 Pro on ARC-AGI-2 benchmark

A new solver for the ARC-AGI-2 visual reasoning benchmark has achieved a top score of 72.9% on the semi-private evaluation set, outperforming leading frontier models like GPT-5.2 Pro and Gemini 3 Pro. The solver employs a modality-driven search strategy, generating reasoning candidates across text, image, and code, and uses a holistic judging approach to compare these candidates within a single prompt. This method effectively identifies correct minority hypotheses, even when the primary modal answer is incorrect. The research also highlights that prescriptive prompting and iterative refinement can reduce hypothesis diversity and degrade performance. AI

IMPACT Sets a new benchmark for visual reasoning, potentially influencing future LLM evaluation and development strategies.

RANK_REASON Research paper detailing a new solver for a benchmark with performance claims against existing models.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New solver beats GPT-5.2 Pro and Gemini 3 Pro on ARC-AGI-2 benchmark

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Research paper detailing a new solver for a benchmark with performance claims against existing models.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
88 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Johan Land ·

    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

    arXiv:2606.31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I prese…

  2. arXiv cs.AI TIER_1 English(EN) · Johan Land ·

    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

    Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a solver for ARC-AGI-2, a few-shot visual rea…