A new benchmark called InteractComp has been developed to evaluate the ability of language agents to handle ambiguous queries during web searches. The benchmark, which includes 210 expert-curated questions, reveals that current agents struggle significantly with ambiguity, with the best-performing model achieving only 13.73% accuracy. The research indicates that enabling agents to interact and ask clarifying questions dramatically improves performance, highlighting a critical gap in their current capabilities despite advancements in general search performance. AI
IMPACT Highlights a critical gap in current AI search agent capabilities, suggesting a need for improved interaction and disambiguation mechanisms.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →