A new recurrent latent reasoning model, significantly smaller than typical transformer models, has achieved a notable score of 29.5% on the ARC-AGI-1 benchmark. This model operates outside the established cost-accuracy frontier for this benchmark, with a reported cost of $0.0007 per task. While promising, further evaluation is needed as the model is scaled to larger parameter sizes. AI
IMPACT Demonstrates potential for smaller, more efficient models to achieve competitive performance on complex reasoning tasks.
RANK_REASON The cluster reports on a new research paper detailing a model's performance on a benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →