PulseAugur
EN
LIVE 08:00:12

New benchmark reveals search agents fail on ambiguous queries

A new benchmark called InteractComp has been developed to evaluate the ability of language agents to handle ambiguous queries during web searches. The benchmark, which includes 210 expert-curated questions, reveals that current agents struggle significantly with ambiguity, with the best-performing model achieving only 13.73% accuracy. The research indicates that enabling agents to interact and ask clarifying questions dramatically improves performance, highlighting a critical gap in their current capabilities despite advancements in general search performance. AI

IMPACT Highlights a critical gap in current AI search agent capabilities, suggesting a need for improved interaction and disambiguation mechanisms.

RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals search agents fail on ambiguous queries

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Mingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong, Jiayi Zhang, Fashen Ren, Jinyi Bai, Fuzhen Yang, Dayi Miao, Zhaoyang Yu, Yifan Wu, Yanfei Zhang, Fengwei Teng, Yingjia Wan, Song Hu, Yude Li, Xin Jin, Conghao Hu, Haoyu Li, Qirui Fu, Tai Zhong, Xinyu Wang, … ·

    InteractComp: Evaluating Search Agents With Ambiguous Queries

    arXiv:2510.24668v2 Announce Type: replace Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a p…