Perplexity has introduced Q2D-Web, a new benchmark designed to evaluate AI agents. This benchmark includes a dataset of 70,000 agent queries, a corpus of 190 million web documents, and three distinct sets of relevance judgments to assess agent performance. AI
IMPACT Provides a new standard for measuring AI agent performance on web-based tasks.
RANK_REASON The item describes the release of a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →