PulseAugur
EN
LIVE 19:31:14

New 'Schema' Harness Boosts ARC-AGI-3 Scores to 99% with Claude Opus 4.8

A new AI harness named "Schema" has been developed, which significantly improves performance on the ARC-AGI-3 benchmark. When used with Anthropic's Claude Opus 4.8 and Meta's Fable 5, Schema achieves a 99% score on the benchmark. A separate test using OpenAI's GPT-5.6 Sol model yielded a 95.35% score. Schema's improvements stem from its novel approach to processing observations, testing predictions, and executing plans, rather than altering the underlying model weights. AI

IMPACT This development could lead to more effective AI agents capable of complex reasoning and problem-solving.

RANK_REASON The item describes a new harness that achieves a high score on a benchmark, which is a research milestone. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/MachineLearning →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New 'Schema' Harness Boosts ARC-AGI-3 Scores to 99% with Claude Opus 4.8

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item describes a new harness that achieves a high score on a benchmark, which is a research milestone. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/we_are_mammals ·

    New Fable5/Opus4.8 harness called "Schema" claims 99% on ARC-3 [R]

    <!-- SC_OFF --><div class="md"><blockquote> <p>Schema, the harness we introduce today, reaches 99% on the ARC‑AGI‑3 Public set using Claude Opus 4.8 and Fable 5, and 95.35% using GPT‑5.6 Sol. It does not change the underlying model weights. Instead, it changes the process around …