PaperBench: Evaluating AI’s Ability to Replicate AI Research
PulseAugur coverage of PaperBench: Evaluating AI’s Ability to Replicate AI Research — every cluster mentioning PaperBench: Evaluating AI’s Ability to Replicate AI Research across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Super Library Agent streamlines multi-application development by consolidating shared code
Researchers have introduced the Super Library Agent, a novel approach to managing portfolios of related software applications. This agent is designed to generate and maintain multiple codebases simultaneously, ensuring …
-
Alibaba releases open-weight Qwen3.8-Max with 2.4T parameters
Alibaba has released Qwen3.8-2.4T-A95B, marking the first open-weight release of a model in its Qwen-Max class. This new model boasts 2.4 trillion total parameters, with 95 billion active parameters per forward pass, ut…
-
Researchers propose ARA format to replace PDF for AI-native scientific papers
A new research artifact format called ARA (Agent-Native Research Artifacts) is proposed as a successor to the traditional PDF for scientific papers. Developed by researchers from multiple institutions, ARA aims to make …
-
Alibaba's Qwen3.8-Max model surpasses GPT-5.6 and Claude Fable 5 on benchmarks
Alibaba's Qwen team has released Qwen3.8-Max, a 2.4-trillion-parameter multimodal model that reportedly outperforms GPT-5.6 Sol Max and Anthropic's Fable 5 on the OSWorld-Verified benchmark. The model is claimed to auto…
-
Alibaba launches Qwen3.8, enhancing coding and office AI capabilities · 2 sources tracked
Alibaba has officially launched its new flagship large language model, Qwen3.8, boasting a total parameter count of 2.4 trillion. This advanced model demonstrates significant improvements in programming and professional…
-
LLM-generated rubrics show bias toward high scores in paper reproduction
A new meta-evaluation of LLM-generated rubrics for paper reproduction reveals that while these rubrics can improve evaluation alignment, they often exhibit biases. The study found that LLM-generated rubrics tend to be o…
-
STORM system improves multi-agent code collaboration with state management
Researchers have introduced STORM, a novel state-oriented management system designed to enhance collaboration among multiple AI agents working on shared codebases. Unlike existing methods that rely on workspace isolatio…