deep research agents
PulseAugur coverage of deep research agents — every cluster mentioning deep research agents across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New DualStake method improves confidence calibration in AI research agents
Researchers have developed DualStake, a novel method to improve the reliability of confidence scores in deep research agents. These agents, used for knowledge-intensive tasks, often exhibit overconfidence, which can und…
-
Deep Research Agents Transform Complex Questions into Trusted Answers
This article explores the concept of deep research agents, which are designed to transform complex questions into reliable answers. It delves into how these agents function and their potential to enhance information ret…
-
New benchmark SciHazard measures LLM scientific safety risks
Researchers have introduced SciHazard, a new benchmark designed to evaluate the safety risks associated with large language models (LLMs) in scientific contexts. The benchmark includes 2400 hazardous questions and 600 o…
-
New WorldCupArena benchmark evaluates AI football forecasting capabilities
A new benchmark called WorldCupArena has been introduced for evaluating language models and deep-research agents on their ability to forecast football (soccer) matches. The benchmark uses the 2026 FIFA World Cup as its …
-
AI scams pose $40B threat; AI nudify apps removed from app stores
This week's cybersecurity highlights include a warning about AI scams potentially costing Americans $40 billion annually, and the discovery that deep-research agents can be compromised through user-generated content. Ad…
-
Research agents use question decomposition for complex business analysis
Deep research agents can effectively tackle complex business questions by first breaking them down into focused, searchable sub-questions. This decomposition allows agents to systematically gather and evaluate informati…
-
New DR-Arena framework automates LLM agent evaluation
Researchers have developed DR-Arena, an automated evaluation framework designed to assess the capabilities of deep research agents, which are advanced large language models capable of autonomous investigation. Unlike st…
-
New research shows 13 words can poison LLMs via user content
A new research paper details a method for poisoning Large Language Models (LLMs) by subtly altering user-generated content. The study suggests that as few as 13 words can be sufficient to compromise the model's integrit…
-
Research Paper Reveals User-Generated Content Can Poison Deep-Research AI Agents
A new research paper details a vulnerability in deep-research agents, which can be compromised through user-generated content. The study, available on arXiv, explores how malicious input can poison these AI systems. Thi…
-
Research paper warns of 'Search-Time Contamination' inflating AI agent benchmarks
A new research paper identifies a problem called Search-Time Contamination (STC) in deep research agents that use web search for evaluation. This contamination occurs when agents retrieve benchmark metadata, question co…
-
New benchmark reveals LLM judges unreliable for research agents
Researchers have developed a new benchmark called REFLECT to evaluate the reliability of Large Language Models (LLMs) when used as judges for deep research agents. These agents automate complex information-seeking tasks…