PulseAugur
EN
LIVE 21:47:49

Perplexity launches Q2D-Web benchmark for retrieval in agentic RAG systems

Perplexity has introduced Q2D-Web, a new benchmark and leaderboard designed to evaluate retrieval performance in agentic RAG systems. The benchmark utilizes a large corpus of 190 million web documents and over 69,000 agent-reformulated queries across ten languages. Initial evaluations show that Perplexity's own pplx-embed-v1-4b model leads in web ranking and combined recall, while Nemotron-3-Embed-8B excels in citation retrieval. AI

IMPACT Establishes a new standard for evaluating retrieval in agentic RAG systems, potentially driving improvements in model performance and benchmark design.

RANK_REASON The cluster describes the release of a new benchmark and leaderboard for evaluating AI systems, which falls under research.

Read on X — Perplexity →

AI-generated summary · Google Gemini · from 9 sources. How we write summaries →

Perplexity launches Q2D-Web benchmark for retrieval in agentic RAG systems

How we ranked this

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster describes the release of a new benchmark and leaderboard for evaluating AI systems, which falls under research.
Source corroboration
9 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [9]

  1. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    To request an evaluation, submit a publicly available Hugging Face retrieval model through our evaluation request form: https://t.co/yGvVSxnrC3

    To request an evaluation, submit a publicly available Hugging Face retrieval model through our evaluation request form: https://t.co/yGvVSxnrC3 The full technical report with methodology and findings is available here: https://t.co/G0AKX2N8r5

  2. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    We evaluated 13 retrieval models across three relevance sets, using Recall@1000 as the primary metric.

    We evaluated 13 retrieval models across three relevance sets, using Recall@1000 as the primary metric. pplx-embed-v1-4b leads Web Ranking (65.73) and Combined (69.11), while Nemotron-3-Embed-8B leads Citation (61.68). https://t.co/z81c25TGMx

  3. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    Reciprocal rank fusion (RRF) subsampling preserves the full-corpus model ranking on Combined Recall@1000 using only 31.7% of the documents.

    Reciprocal rank fusion (RRF) subsampling preserves the full-corpus model ranking on Combined Recall@1000 using only 31.7% of the documents. For pplx-embed-v1-4b, this reduces evaluation from 4,608 to roughly 1,500 H200 GPU-hours. https://t.co/ZPTvJ1xYoN

  4. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    Q2D-Web uses agent citations, production web rankings, and a combined set expanded with LLM judgments.

    Q2D-Web uses agent citations, production web rankings, and a combined set expanded with LLM judgments. These three relevance sets reduce false negatives and reliance on a single labeling pipeline, while testing how relevance definitions affect model performance. https://t.co/g6m…

  5. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    The corpus combines the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH.

    The corpus combines the top 5,000 production retrieval results per query, deduplicated with MinHash-LSH. Each document is a plausible match for at least one query, including difficult distractors that match the topic but miss a required date, entity, or version. https://t.co/6XK…

  6. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    Q2D-Web covers programming, law, health, science, and finance alongside consumer goods, travel, entertainment, and local information. Queries span ten languages

    Q2D-Web covers programming, law, health, science, and finance alongside consumer goods, travel, entertainment, and local information. Queries span ten languages, with English accounting for 65.8%. https://t.co/GdUEviYaGu

  7. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    Q2D-Web draws on 23,000 PII-free production searches across ten languages and dozens of domains, collected over nine months.

    Q2D-Web draws on 23,000 PII-free production searches across ten languages and dozens of domains, collected over nine months. Agents reformulate user requests into primary and support queries, each evaluated independently with its own relevance judgments. https://t.co/xUI4OxiBcr

  8. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    Realistic retrieval evaluation requires large corpora and query sets, with deep relevance judgments to reduce false negatives.

    Realistic retrieval evaluation requires large corpora and query sets, with deep relevance judgments to reduce false negatives. Q2D-Web combines 190M web documents, 69,721 agent-reformulated queries, and 99.6 positive judgments per query on average in its combined set. https://t.…

  9. X — Perplexity TIER_1 English(EN) · perplexity_ai ·

    We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems.

    We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems. Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries. Read more: https://t.co/s476SxkE1L