English(EN)Q2D-Web covers programming, law, health, science, and finance alongside consumer goods, travel, entertainment, and local information. Queries span ten languages
To request an evaluation, submit a publicly available Hugging Face retrieval model through our evaluation request form: https://t.co/yGvVSxnrC3
The full technical report with methodology and findings is available here:
https://t.co/G0AKX2N8r5
We evaluated 13 retrieval models across three relevance sets, using Recall@1000 as the primary metric.
pplx-embed-v1-4b leads Web Ranking (65.73) and Combined (69.11), while Nemotron-3-Embed-8B leads Citation (61.68). https://t.co/z81c25TGMx
Reciprocal rank fusion (RRF) subsampling preserves the full-corpus model ranking on Combined Recall@1000 using only 31.7% of the documents.
For pplx-embed-v1-4b, this reduces evaluation from 4,608 to roughly 1,500 H200 GPU-hours. https://t.co/ZPTvJ1xYoN
Q2D-Web uses agent citations, production web rankings, and a combined set expanded with LLM judgments.
These three relevance sets reduce false negatives and reliance on a single labeling pipeline, while testing how relevance definitions affect model performance. https://t.co/g6m…
The corpus combines the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH.
Each document is a plausible match for at least one query, including difficult distractors that match the topic but miss a required date, entity, or version. https://t.co/6XK…
Q2D-Web covers programming, law, health, science, and finance alongside consumer goods, travel, entertainment, and local information. Queries span ten languages, with English accounting for 65.8%. https://t.co/GdUEviYaGu
Q2D-Web draws on 23,000 PII-free production searches across ten languages and dozens of domains, collected over nine months.
Agents reformulate user requests into primary and support queries, each evaluated independently with its own relevance judgments. https://t.co/xUI4OxiBcr
Realistic retrieval evaluation requires large corpora and query sets, with deep relevance judgments to reduce false negatives.
Q2D-Web combines 190M web documents, 69,721 agent-reformulated queries, and 99.6 positive judgments per query on average in its combined set. https://t.…
We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems.
Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries.
Read more: https://t.co/s476SxkE1L