DisCoBench
PulseAugur coverage of DisCoBench — every cluster mentioning DisCoBench across labs, papers, and developer communities, ranked by signal.
-
AI agents fail research tasks due to lack of clarifying questions; Hollywood faces copyright issues with Seedance
Recent tests using the DisCoBench benchmark indicate that AI agents struggle with research tasks due to an inability to ask clarifying questions, leading to only a 40% success rate. Separately, the Motion Picture Associ…
-
AI search agents fail by not asking clarifying questions, new benchmark shows
AI search agents struggle not with finding information, but with clarifying ambiguous queries. A new benchmark, DiscoBench, reveals that agents which repeatedly search instead of asking follow-up questions perform worse…
-
DiscoGen system generates billions of ML algorithm discovery tasks
Researchers have introduced DiscoGen, a novel system designed to procedurally generate a vast array of machine learning algorithm discovery tasks. This tool aims to overcome limitations in current task suites, such as p…
-
New benchmark DiscoBench evaluates LLM search agents' ability to handle ambiguous queries
Researchers have introduced DiscoBench, a new benchmark designed to evaluate the ability of large language model (LLM) powered search agents to handle ambiguous queries. The benchmark includes 211 samples and 463 ambigu…
-
New benchmark DiscoBench evaluates LLM search agents' clarification skills
A new benchmark called DiscoBench has been introduced to evaluate the ability of search agents, particularly those powered by LLMs, to handle ambiguous queries. DiscoBench assesses how effectively these agents can ident…