English(EN)DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
新的AI代理处理深度研究和误导性网络数据 · 跟踪4个来源
作者PulseAugur 编辑部·[4 个来源]·
研究人员推出了一系列新的递归自改进代理AREX,用于深度研究任务。AREX在研究和自我改进循环之间交替进行,使用自主上下文更新工具来管理不断增长的交互历史。这种方法使AREX在BrowseComp和Humanity's Last Exam等基准测试中表现优于同等规模的基线。同时,另一项研究引入了DRNOISE,这是一个旨在评估深度研究代理在开放网络上处理误导性信息的能力的基准,突出了存在此类文件时准确性显著下降的问题。
AI
arXiv:2607.21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery…
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a researc…
arXiv cs.CL
TIER_1English(EN)·Jun Nie, Zhiqin Yang, Zhenheng Tang, Yonggang Zhang, Xiaowen Chu, Xinmei Tian, Bo Han·
arXiv:2607.17291v1 Announce Type: cross Abstract: Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents prese…
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-lo…