Anthropic has released Claude Opus 5.5, which is performing exceptionally well on vision tasks and leading SimpleBench. This new model is also showing strong reasoning capabilities on the Terminal-Bench-Science benchmark, rivaling GPT-6 Astra. Meanwhile, OpenAI's GPT-6 family, including Astra, Sol, and Luna, has demonstrated impressive performance in various domains such as NetHack, Code Arena, and DOOM agent tasks. Other notable releases include Gemini 3.8 Flash with a large context window and Xiaomi's omni-modal MiMo-V2.6-Pro, which is open-source and cost-effective. AI
IMPACT Sets new SOTA on vision and reasoning benchmarks, intensifying competition among top AI labs.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Anthropic
- Claude Code
- Codex
- Databricks
- Gemini 3.8 Flash
- Google Cloud Platform
- GPT-6
- GPT-6 Astra
- GPT-6 Luna
- GPT-6 Sol
- Meta*
- Muse Spark 1.3
- Opus 5.5
- SimpleBench
- Terminal-Bench-Science
- Xiaomi MiMo-V2.6-Pro
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →