ARC-AGI-1
PulseAugur coverage of ARC-AGI-1 — every cluster mentioning ARC-AGI-1 across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New 'Looped Flows' method boosts AI reasoning on complex tasks
Researchers have introduced "looped flows," a novel approach to training recurrent neural networks that enhances their ability to solve complex problems by allowing for more computational updates during inference. This …
-
Astra language model achieves high scores on ARC-AGI benchmark without CoT
The Astra language model has achieved impressive scores on the ARC-AGI benchmark, reaching 97% on ARC-AGI-3 and 86% on ARC-AGI-1. Notably, these high scores were attained without the use of Chain-of-Thought (CoT) prompt…
-
Small Transformer Achieves 44% on ARC-AGI-1 Benchmark for 67 Cents
A researcher has developed a small transformer model that achieves 44% accuracy on the ARC-AGI-1 benchmark, a significant feat accomplished in just 1.5 hours and costing only 67 cents. This model surpasses many existing…
-
150M param recurrent model achieves strong ARC-AGI-1 score at low cost · 2 sources tracked
A new recurrent latent reasoning model, significantly smaller than typical transformer models, has achieved a notable score of 29.5% on the ARC-AGI-1 benchmark. This model operates at a very low cost of $0.0007 per task…
-
Pathway's BDH-CQ model sets new ARC-AGI-1 cost-efficiency frontier
Pathway has announced BDH-CQ, a 150 million parameter post-Transformer model that achieves a score of 29.5% on the ARC-AGI-1 benchmark. This model reportedly sets a new cost-efficiency frontier, with a computed cost of …
-
DeepSeek V4 Flash 0731 achieves high scores on ARC-AGI benchmark
DeepSeek has released its V4 Flash 0731 model, featuring three reasoning settings: Low, High, and Max. The model achieved impressive scores on the ARC Prize benchmark, a test for abstract reasoning in AI systems. Notabl…
-
BDH-CQ model achieves new state-of-the-art in AI reasoning cost efficiency
Researchers have developed BDH-CQ, a novel reasoning model that integrates in-context learning with recurrent latent reasoning. This model updates its memory with inference-time inputs and iteratively computes solutions…
-
TraceViT model enhances AI visual abstract reasoning with step-by-step supervision
Researchers have introduced TraceViT, a novel looped visual reasoner designed to improve performance on abstract reasoning tasks. Unlike previous methods that only constrain the final output, TraceViT is trained to foll…
-
AI benchmark scores predictable from just two factors, study finds
A new research paper proposes a method called BenchPress that can predict a frontier model's performance across numerous benchmarks using only two key scores. The study analyzed 84 models and 133 benchmarks, finding tha…