Perplexity has introduced WANDR, a new open benchmark designed to evaluate research agents. This benchmark comprises 500 tasks that require agents to find and cite evidence to support their discoveries. In initial tests, Perplexity Search as Code achieved the highest score with a 0.363 soft F1. AI
IMPACT This benchmark could accelerate the development and evaluation of AI agents capable of complex information retrieval and evidence-based reasoning.
RANK_REASON The cluster describes the release of a new benchmark for evaluating AI research agents. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →