PulseAugur
实时 09:33:12
English(EN) DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

新基准显示深度研究代理未能达到专家标准

研究人员推出了 DeepResearch Bench II,这是一个旨在评估深度研究代理 (DRAs) 能力的新基准。该基准包含 22 个领域的 132 项研究任务,每项任务要求代理生成一份报告,并根据 9,430 个细粒度评分标准进行评估。这些评分标准源自专家撰写的文章,并通过人类-LLM 管道进行完善,侧重于信息回忆、分析和呈现。初步评估显示,即使是先进的 DRAs 也未能满足超过 50% 的标准,表明与人类研究能力相比存在显著差距。该基准、评估脚本和评分标准已公开发布,以鼓励该领域的进一步发展。 AI

影响 该基准将推动 AI 代理在执行和报告复杂研究任务方面的能力改进。

排序理由 该集群描述了一个用于评估 AI 代理的新学术基准,包括其方法和发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准显示深度研究代理未能达到专家标准

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个用于评估 AI 代理的新学术基准,包括其方法和发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao ·

    DeepResearch Bench II:通过专家报告中的标准诊断深度研究代理

    arXiv:2601.08536v3 Announce Type: replace Abstract: Deep Research Agents (DRA) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. Prior benchmarks often either under-evaluate a system's ability to produce meaningful insights…