Claude Opus 5 has achieved the second-highest score on the SimpleBench benchmark, trailing only Fable 5 by a narrow margin. The benchmark evaluates models on spatio-temporal reasoning, social intelligence, and their ability to handle trick questions. Opus 5 significantly outperformed previous versions like Opus 4.6, 4.7, and 4.8, and all other models tested on this specific evaluation. AI
IMPACT Sets a new performance benchmark for reasoning and robustness, indicating progress in advanced AI capabilities.
RANK_REASON The cluster reports on a specific benchmark performance for an AI model, which falls under research.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →