PulseAugur
EN
LIVE 19:05:57
ENTITY research agents

research agents

PulseAugur coverage of research agents — every cluster mentioning research agents across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
1 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 4 TOTAL
  1. TOOL · CL_206356 ·

    New WANDR benchmark tests AI agents on deep data collection tasks

    A new benchmark called WANDR has been introduced for evaluating the capabilities of research agents in wide and deep data collection tasks. WANDR comprises 500 realistic scenarios that require agents to discover a broad…

  2. COMMENTARY · CL_166289 ·

    22 common failure modes identified in LLM agents

    LLM agents, regardless of their specialization like coding or research, exhibit 22 consistent failure modes rather than unique bugs. These failures can be categorized, and specific prompts can mitigate them. The effecti…

  3. TOOL · CL_150809 ·

    Perplexity launches WANDR benchmark for AI research agents

    Perplexity has introduced WANDR, a new open benchmark designed to evaluate research agents. This benchmark comprises 500 tasks that require agents to find and cite evidence to support their discoveries. In initial tests…

  4. RESEARCH · CL_82130 ·

    LLM research agents show low overfitting due to strategy compressibility

    Researchers have investigated why machine learning, particularly when driven by large language models (LLMs), exhibits surprisingly little overfitting despite adaptive benchmark use. Their study on LLM-driven research a…