BrowseComp+
PulseAugur coverage of BrowseComp+ — every cluster mentioning BrowseComp+ across labs, papers, and developer communities, ranked by signal.
- instance of CatalyzeX 70%
- used by Grpo 70%
- used by alphaXiv 70%
- used by ScienceCast 70%
- instance of DeepSearchQA 70%
- instance of Gotit.pub 70%
- used by BrowseComp-ZH 70%
- used by DagsHub 70%
- instance of React 70%
- instance of Generative Ai Interactive Agents 70%
- used by Deep Research 70%
- used by Generative Ai Interactive Agents 60%
7 day(s) with sentiment data
-
New DRBENCHER benchmark tests AI agents' combined browsing and math skills
Researchers have introduced DRBENCHER, a new benchmark designed to evaluate AI agents' ability to combine web browsing with multi-step mathematical computations. Unlike previous benchmarks that assess these skills in is…
-
New frameworks enhance AI agent reasoning and grounding capabilities
Researchers are developing new frameworks to improve the capabilities of AI agents, particularly in their ability to perform long-horizon reasoning and reflection. LoongReflect focuses on memory control and combines glo…
-
ChatGPT's Work agent mode shows mixed results on complex tasks
A recent 20-hour test of ChatGPT's agent mode, now called Work and powered by GPT-5.6, revealed mixed results. While it shows promise in tasks like data aggregation and personalized messaging, it struggles with complex …
-
New research aims to improve retrieval-augmented search agents · 2 sources tracked
Two new research papers propose methods to improve the efficiency and effectiveness of retrieval-augmented search agents. The first paper, "HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents," intro…
-
New CRISP framework trains LLM search agents to be more efficient
Researchers have introduced CRISP, a new framework designed to train more efficient deep search agents powered by large language models. Unlike previous methods that simply reduce tool usage, CRISP identifies and preser…
-
New credit assignment methods enhance AI search agent training · 3 sources tracked
Researchers have developed new methods for training long-horizon search agents, which are AI systems designed to perform complex, multi-step tasks. One approach, ABSeeker, uses Answer-Backtracked Credit Assignment (ABC)…
-
Anthropic's Claude Opus 5 excels in bio/cyber tasks, bypassing Fable 5 restrictions
Anthropic has released Claude Opus 5, positioning it as a powerful tool for computational biology and cybersecurity tasks. While the "frontier" model, Fable 5, is heavily restricted in these domains, and Mythos 5 is gat…
-
Anthropic's Claude Cookbook offers advanced agent-building recipes · 2 sources tracked
Anthropic has released a collection of resources and tutorials, dubbed the Claude Cookbook, detailing how to build and deploy advanced AI agents using its Claude models and SDKs. These resources cover a range of applica…
-
New AI agents tackle deep research and misleading web data · 4 sources tracked
Researchers have introduced AREX, a new family of recursively self-improving agents designed for deep research tasks. AREX alternates between research and self-improvement loops, using an autonomous context-update tool …
-
Agents-A1-4B model shows strong performance in long-horizon search
Agents-A1-4B, a new model developed by InternScience, demonstrates strong performance across various benchmarks, particularly in long-horizon search and agentic tasks. The model, which is based on Qwen3.7-4B, significan…
-
New STAMP method improves credit assignment for deep search agents
Researchers have introduced STAMP, a novel method for improving credit assignment in deep search agents. This approach addresses the 'reward-credit mismatch' by providing targeted credit to actions that expose supportin…
-
New framework enables AI agents to self-improve in verifiable web environments
Researchers have introduced DeepSearch-Evolve, a self-distillation framework designed to train web agents within the DeepSearch-World environment. This framework aims to overcome challenges in agent training by enabling…
-
DeepSearch-Evolve framework trains web agents via self-distillation in verifiable environment
Researchers have introduced DeepSearch-Evolve, a self-distillation framework designed to train web agents. This framework utilizes DeepSearch-World, a verifiable environment containing 420,000 multi-hop question-answeri…
-
New method mines agent skills from interaction data, but policy improvement is limited
Researchers have developed a method to automatically generate skill libraries for computer-using agents by mining interaction trajectories. The process involves segmenting graphical user interface (GUI) trajectories, cl…
-
New LLM Training Methods Optimize Data Scheduling for Efficiency and Performance
Researchers have developed new methods for optimizing the training of large language models (LLMs) through advanced data scheduling techniques. One approach, the Holistic Data Scheduler (HDS), uses multi-objective reinf…
-
Perplexity Integrates Deep Research with Multi-Model Orchestration System
Perplexity has integrated its Deep Research feature into its Computer orchestration system, enhancing its ability to break down complex questions into subtasks. These subtasks are then routed across more than 20 differe…
-
TreeSeeker framework enhances AI deep search with controlled trial-and-error
Researchers have introduced TreeSeeker, a novel framework designed to improve the efficiency of deep search agents. This system structures search processes as a tree, allowing agents to explore multiple potential paths …
-
New Korean web-browsing benchmark reveals LLM performance gaps
Researchers have introduced K-BrowseComp, a new benchmark designed to evaluate the web-browsing agent capabilities of large language models specifically within Korean contexts. The benchmark comprises 400 problems, with…
-
Author warns AI evaluations are unreliable, risking unseen harms
The author argues that current AI evaluation methods are unreliable and systematically flawed, posing significant risks. They highlight issues like models gaming evaluations, distribution shifts rendering metrics inaccu…
-
New benchmark LiveBrowseComp tests LLM search agents' true discovery skills
A new research paper introduces LiveBrowseComp, a benchmark designed to assess whether large language model (LLM) search agents truly discover new information or merely verify their existing internal knowledge. The stud…