Keenable AI has open-sourced NEEDLE, a new benchmark designed to evaluate web search APIs by dynamically generating query sets hourly and daily. This approach prevents agents from accessing pre-existing answers, ensuring a more accurate assessment of their retrieval capabilities. NEEDLE categorizes queries into News, Finance, Scholar, Deep-tail, and Legal, with scoring methods tailored to each vertical. The benchmark also establishes an 'ultimate ceiling' by pooling results from all tested search APIs, allowing for a clear distinction between ranking issues and market-wide retrieval problems. AI
IMPACT Provides a more robust method for evaluating search agents, potentially driving improvements in retrieval and ranking quality across the industry.
RANK_REASON Open-source release of a new benchmark for evaluating AI search capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CourtListener.com
- DeepResearchGym
- Europe PubMed Central
- Global Legal Entity Identifier Foundation
- Google Trends
- Keenable AI
- LRAT
- OpenResearcher
- OpenRouter
- Python
- RSS
- SEC XBRL
- Wikidata
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →