PulseAugur
实时 23:02:48
English(EN) Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

新的 Q2D-Web 基准测试评估 Agentic RAG 系统的检索能力

研究人员推出了 Q2D-Web,这是一个新的大规模基准测试,旨在评估 Agentic 检索增强生成 (RAG) 管道中的检索系统。该基准测试通过将包含 1.9 亿个文档的大型网络语料库与源自真实用户对话的 7 万个 Agent 重新制定的搜索查询配对,解决了现有数据集的局限性。Q2D-Web 提供了多个相关性判断集,并探索了子语料库采样技术,以在保持模型排名准确性的同时,加快评估速度。 AI

影响 为评估复杂 RAG 系统中的检索组件提供了新的标准,有望提高 Agentic AI 的性能。

排序理由 该集群描述了一个用于评估 AI 检索系统的新学术基准测试。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的 Q2D-Web 基准测试评估 Agentic RAG 系统的检索能力

本文如何被排名

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一个用于评估 AI 检索系统的新学术基准测试。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Denis Bykov ·

    Q2D-Web: 智能体式 RAG 系统中检索的大规模基准测试

    Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query.…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Q2D-Web: 智能体式 RAG 系统检索的大规模基准测试

    Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query.…