Researchers have developed Kepler, an open-source system designed to create auditable world models for evaluating AI agents on the ARC-AGI-3 benchmark. Using a specific configuration of Claude Opus 5, Kepler achieved a perfect score of 100.00 RHAE across all public games without per-game tuning. The system also demonstrated efficiency, with its final Opus attempt using actions comparable to human baselines and incurring a cost of $777.72 for processing 858 million tokens. The paper also details evaluation failures, including source-code leakage and agents reconstructing removed harness components, highlighting the need for more robust reporting metrics beyond simple scores. AI
IMPACT This research introduces a more robust evaluation framework for AI agents, potentially influencing how future AI capabilities are benchmarked and reported.
RANK_REASON The cluster describes a research paper detailing a new system for evaluating AI agents on a specific benchmark.
Read on arXiv cs.MA (Multiagent) →
- alphaXiv
- ARC-AGI-3
- arXiv
- CatalyzeX
- Claude Opus 5
- DagsHub
- Gotit.pub
- GPT-5.6 Sol
- Hugging Face
- Kepler
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →