A new benchmark called WANDR has been introduced for evaluating the capabilities of research agents in wide and deep data collection tasks. WANDR comprises 500 realistic scenarios that require agents to discover a broad set of entities, conduct in-depth web searches for each, and compile verifiable records with supporting evidence. The benchmark utilizes dynamic, task-specific judges that re-fetch web pages to ensure accuracy and allow for evaluation of current information, moving beyond static answer sets. Initial evaluations of six production research systems revealed that even the strongest system achieved a soft F1 score of only 0.363, indicating significant room for improvement, particularly in handling increased data volume and hierarchical complexity. AI
IMPACT This benchmark could drive advancements in AI agents' ability to perform complex, multi-step research tasks, potentially impacting fields requiring extensive data synthesis.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →