A new benchmark called GoBench, designed to measure general reasoning abilities through the game of Go, has been introduced. The benchmark shows that GPT-6 Astra achieved the highest Elo rating at 2568, surpassing other models like Sol (1929 Elo) and Opus 5 high (2076 Elo). Furthermore, when integrated with coding capabilities, GPT-6 Astra's Codex variant reached an Elo of 3563, significantly outperforming Codex with Sol's 2656 Elo. AI
IMPACT Establishes a new evaluation method for AI reasoning, potentially driving development towards more general intelligence.
RANK_REASON Introduction of a new benchmark and performance results for AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →