PulseAugur
实时 10:01:02
English(EN) Do Web Agents Investigate Before They Decide?

新的MIRAGE基准测试了网络代理的调查能力

一个名为MIRAGE的新基准已被开发出来,用于评估自主网络代理的调查能力。该基准包含750个跨维基百科取证、购物管理和Reddit版主任务的多步决策任务,每个任务都有一个可能具有误导性的可见表面上下文和一个需要主动调查的隐藏上下文。对八个LLM代理的评估显示,尽管代理可以访问相关页面,但它们难以提取决定性证据,在面对矛盾信息时常常失败,并表现出显著的调查性幻觉率。 AI

影响 突出了当前AI代理能力的一个关键差距,表明需要改进证据检索和不确定性下的推理能力。

排序理由 该集群包含一篇介绍用于评估AI代理的新基准的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的MIRAGE基准测试了网络代理的调查能力

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇介绍用于评估AI代理的新基准的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Syed Nazmus Sakib, Nafiul Haque, Tapodhir Karmakar Taton, Shahrear Bin Amin, Shifat E. Arman ·

    网络代理在做决定前会调查吗?

    arXiv:2602.05354v3 Announce Type: replace Abstract: Autonomous web agents are increasingly deployed in moderation and policy enforcement, where correct decisions often depend on evidence that is not immediately visible and must be actively investigated. Yet existing benchmarks la…