OpenAI claims its new GPT-5.6 Sol model can outperform Anthropic's Opus 5 on the ARC-AGI-3 benchmark. However, this superior score of 38.3% was achieved using OpenAI's proprietary API features, including retained reasoning and context compaction. When tested in the official, provider-neutral ARC-AGI-3 environment, GPT-5.6 Sol scored only 7.8%, while Opus 5 achieved 30.2% without such specialized settings. AI
IMPACT Highlights ongoing challenges in fair and standardized benchmarking of large language models.
RANK_REASON The cluster reports on claims made by OpenAI about a model's performance on a benchmark, but focuses on the controversy surrounding the testing methodology rather than an official release or research paper.
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →