The Astra language model has achieved impressive scores on the ARC-AGI benchmark, reaching 97% on ARC-AGI-3 and 86% on ARC-AGI-1. Notably, these high scores were attained without the use of Chain-of-Thought (CoT) prompting, suggesting a strong inherent reasoning capability within the model. AI
IMPACT Demonstrates strong reasoning capabilities in language models, potentially influencing future benchmark development and model training strategies.
RANK_REASON Research benchmark results for a language model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →