research agents
PulseAugur coverage of research agents — every cluster mentioning research agents across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
22 common failure modes identified in LLM agents
LLM agents, regardless of their specialization like coding or research, exhibit 22 consistent failure modes rather than unique bugs. These failures can be categorized, and specific prompts can mitigate them. The effecti…
-
Perplexity launches WANDR benchmark for AI research agents
Perplexity has introduced WANDR, a new open benchmark designed to evaluate research agents. This benchmark comprises 500 tasks that require agents to find and cite evidence to support their discoveries. In initial tests…
-
LLM research agents show low overfitting due to strategy compressibility
Researchers have investigated why machine learning, particularly when driven by large language models (LLMs), exhibits surprisingly little overfitting despite adaptive benchmark use. Their study on LLM-driven research a…