GPT-6 Astra has successfully solved the final remaining problem in the FrontierMath Tier 4 benchmark, a set of research-level mathematical problems designed to challenge advanced AI models. This achievement marks a significant milestone, as all problems within this rigorous test suite have now been solved by AI at least once. While Astra's direct score was 97.6%, its breakthrough on the previously unsolved problem signifies that the FrontierMath Tier 4 benchmark is now considered saturated, with AI demonstrating advanced mathematical reasoning capabilities. AI
IMPACT Sets a new benchmark for AI mathematical reasoning, potentially accelerating research into AI's ability to solve complex, unsolved mathematical problems.
RANK_REASON The cluster reports on a new model, GPT-6 Astra, achieving a significant milestone on a research-level benchmark (FrontierMath Tier 4), which is a direct disclosure from a frontier lab. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Claude Fable-5
- Claude Opus 4.7
- Epoch AI
- FrontierMath Tier 4
- GPT-5.5
- GPT-5.6 Sol
- GPT-6 Astra
- GSM8K
- Jay Pantone
- Lean
- Marquette University
- mathematics-dataset
- OpenAI
- Richard Borcherds
- Terence Tao
- Timothy Gowers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →