A new AI model, GPT-6 Astra, has reportedly achieved a score of 98.6% on the ARC-AGI-3 benchmark, significantly surpassing the previous GPT-5.6 "Sol" model's 7.8%. Despite this high performance, OpenAI cautions against direct comparisons due to variations in evaluation methodologies. While GPT-6 Astra demonstrates advanced problem-solving capabilities and has aided in mathematical research, its reasoning consistency is noted as inconsistent, and achieving AGI status remains unconfirmed. AI
IMPACT Sets a new benchmark for AI reasoning capabilities, potentially influencing future model development and AGI research.
RANK_REASON The item reports on a new benchmark score for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →