Anthropic's Claude Opus 5 has achieved a score of 30.2% on the ARC-AGI 3 benchmark. This performance metric was shared across several online communities, highlighting the model's capabilities in complex reasoning tasks. The benchmark is known for its difficulty, making this score a notable data point for the model's development. AI
IMPACT Demonstrates progress in complex reasoning capabilities for large language models.
RANK_REASON Research benchmark result for an AI model.
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →