A new benchmark called DiscoBench has been introduced to evaluate the ability of search agents, particularly those powered by LLMs, to handle ambiguous queries. DiscoBench assesses how effectively these agents can identify ambiguity, ask clarifying questions, and recover correct reasoning paths through user interaction. Experiments revealed that agents often struggle with ambiguity detection and that repeatedly searching can be less effective than asking for clarification or even guessing, highlighting a gap in current agents' interactive problem-solving capabilities. AI
IMPACT Highlights a critical gap in LLM agent capabilities for interactive problem-solving and ambiguity resolution.
RANK_REASON The item describes a new academic benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- Clarification Questioning
- DirectGuess
- DiscoBench
- Google Deep Search
- Information-Seeking Tasks
- LLMs
- Search Agents
- SearchHeavyGuess
- SearchThenAsk
- User Simulator
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →