PulseAugur
实时 15:02:05
English(EN) DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

新的AI代理处理深度研究和误导性网络数据 · 跟踪4个来源

研究人员推出了一系列新的递归自改进代理AREX,用于深度研究任务。AREX在研究和自我改进循环之间交替进行,使用自主上下文更新工具来管理不断增长的交互历史。这种方法使AREX在BrowseComp和Humanity's Last Exam等基准测试中表现优于同等规模的基线。同时,另一项研究引入了DRNOISE,这是一个旨在评估深度研究代理在开放网络上处理误导性信息的能力的基准,突出了存在此类文件时准确性显著下降的问题。 AI

影响 这些进展凸显了AI代理在复杂研究方面的能力进步及其在应对误导性信息方面的鲁棒性。

排序理由 该集群包含两篇介绍新AI代理架构和基准的研究论文。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新的AI代理处理深度研究和误导性网络数据 · 跟踪4个来源

报道来源 [4]

  1. arXiv cs.AI TIER_1 English(EN) · Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zh… ·

    AREX:迈向深度研究的递归自改进智能体

    arXiv:2607.21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    AREX:迈向深度研究的递归自改进智能体

    Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a researc…

  3. arXiv cs.CL TIER_1 English(EN) · Jun Nie, Zhiqin Yang, Zhenheng Tang, Yonggang Zhang, Xiaowen Chu, Xinmei Tian, Bo Han ·

    DRNOISE:在误导性证据环境中对深度研究代理进行基准测试

    arXiv:2607.17291v1 Announce Type: cross Abstract: Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents prese…

  4. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Bo Han ·

    DRNOISE:在误导性证据环境中对深度研究代理进行基准测试

    Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-lo…