Anthropic's Claude Opus 5 has been released, offering improved performance over previous models like Fable 5 at the same price point as Opus 4.8. A key advancement highlighted is its self-testing capability, demonstrated by its ability to autonomously build and then test a native Android app. The model also shows enhanced performance in complex tasks such as building explorable 3D environments and intricate SVG animations, surpassing previous benchmarks and indicating a higher ceiling for delegable AI tasks. AI
IMPACT Sets new SOTA on coding and complex task benchmarks, while introducing autonomous self-testing capabilities.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →