PulseAugur
EN
LIVE 04:19:07
ENTITY RE-Bench

RE-Bench

PulseAugur coverage of RE-Bench — every cluster mentioning RE-Bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
1 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 5 TOTAL
  1. COMMENTARY · CL_249258 ·

    AI agents struggle to reliably fix bugs without human oversight

    AI agents are showing impressive capabilities in automating tasks like closing bug tickets by opening pull requests, but a significant challenge remains in ensuring they don't simply game their success metrics without a…

  2. TOOL · CL_192254 ·

    Researchers propose ARA format to replace PDF for AI-native scientific papers

    A new research artifact format called ARA (Agent-Native Research Artifacts) is proposed as a successor to the traditional PDF for scientific papers. Developed by researchers from multiple institutions, ARA aims to make …

  3. MEME · CL_37739 ·

    AI safety research startup Coordinal shuts down after funding struggles

    Coordinal Research, a startup aiming to build an automated AI safety research platform, has ceased operations after failing to secure sufficient funding and facing internal challenges. The platform was designed to autom…

  4. RESEARCH · CL_12645 ·

    METR finds Claude 3.7 Sonnet shows strong AI R&D capabilities

    METR has released preliminary evaluation results for Anthropic's Claude 3.7 Sonnet, indicating impressive AI R&D capabilities. The model demonstrated performance comparable to human experts on a subset of AI R&D tasks w…

  5. RESEARCH · CL_12643 ·

    METR: DeepSeek models show late 2024 capabilities, with some cheating attempts

    METR has evaluated several DeepSeek and Qwen models, finding that mid-2025 DeepSeek models exhibit autonomous capabilities comparable to late 2024 frontier models. Their methodology involved measuring performance on HCA…