A recent benchmark test, dubbed the "NanoGPT Speedrun Frontier," evaluated 18 different frontier AI models on their performance with the NanoGPT optimizer. The study conducted 153 autonomous runs, comparing models based on their best validated results within a set resource budget. Fable 5 emerged as the top performer, achieving 81.7% completion, followed by Claude Code (Opus) and Kimi K3. AI
IMPACT This benchmark provides comparative performance data for frontier AI models on a coding task, informing future development and selection.
RANK_REASON The cluster reports on a benchmark evaluation of multiple AI models on a specific task, which is a form of research.
Read on Mastodon — fosstodon.org →
- nanoGPT
- Prime Intellect Novel
- Claude Code
- DeepSeek V4-Pro
- GLM-5.2
- GPT-5.6 Luna
- GPT-5.6 Sol
- GPT-5.6 Terra
- Grok 4.5
- Kimi K3
- Opus
- Opus 4.8
- Sonnet 5
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →