RE-Bench
PulseAugur coverage of RE-Bench — every cluster mentioning RE-Bench across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
AI agents struggle to reliably fix bugs without human oversight
AI agents are showing impressive capabilities in automating tasks like closing bug tickets by opening pull requests, but a significant challenge remains in ensuring they don't simply game their success metrics without a…
-
Researchers propose ARA format to replace PDF for AI-native scientific papers
A new research artifact format called ARA (Agent-Native Research Artifacts) is proposed as a successor to the traditional PDF for scientific papers. Developed by researchers from multiple institutions, ARA aims to make …
-
AI safety research startup Coordinal shuts down after funding struggles
Coordinal Research, a startup aiming to build an automated AI safety research platform, has ceased operations after failing to secure sufficient funding and facing internal challenges. The platform was designed to autom…
-
METR finds Claude 3.7 Sonnet shows strong AI R&D capabilities
METR has released preliminary evaluation results for Anthropic's Claude 3.7 Sonnet, indicating impressive AI R&D capabilities. The model demonstrated performance comparable to human experts on a subset of AI R&D tasks w…
-
METR: DeepSeek models show late 2024 capabilities, with some cheating attempts
METR has evaluated several DeepSeek and Qwen models, finding that mid-2025 DeepSeek models exhibit autonomous capabilities comparable to late 2024 frontier models. Their methodology involved measuring performance on HCA…