PulseAugur
中
实时 18:39:51
English(EN) WANDR: A Benchmark for Wide and Deep Research

新的 WANDR 基准测试用于测试 AI 代理在深度数据收集任务中的能力

引入了一个名为 WANDR 的新基准测试,用于评估研究代理在宽泛且深入的数据收集任务中的能力。WANDR 包含 500 个现实场景,要求代理发现广泛的实体集,为每个实体进行深入的网络搜索,并汇编带有支持证据的可验证记录。该基准测试使用动态的、特定任务的裁判,它们会重新抓取网页以确保准确性,并允许评估当前信息,超越了静态答案集。对六个生产研究系统的初步评估显示,即使是最强的系统,软 F1 分数也仅为 0.363,表明在处理增加的数据量和层次复杂性方面仍有很大的改进空间。 AI

影响 该基准测试有望推动 AI 代理执行复杂、多步骤研究任务的能力的进步,可能影响需要广泛数据综合的领域。

排序理由 该集群描述了一个用于评估 AI 代理的新学术基准测试。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 WANDR 基准测试用于测试 AI 代理在深度数据收集任务中的能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个用于评估 AI 代理的新学术基准测试。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
51 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma ·

    WANDR:面向广泛和深度研究的基准测试

    arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), invest…