BrowseComp+
PulseAugur coverage of BrowseComp+ — every cluster mentioning BrowseComp+ across labs, papers, and developer communities, ranked by signal.
- instance of Gotit.pub 70%
- used by DagsHub 70%
- used by alphaXiv 70%
- used by ScienceCast 70%
- instance of CatalyzeX 70%
- used by Gotit.pub 70%
- instance of BrowseComp-ZH 70%
- instance of Generative Ai Interactive Agents 70%
- used by Grpo 70%
- used by BrowseComp-ZH 70%
- instance of React 70%
- used by Deep Research 70%
3 day(s) with sentiment data
-
7B model ZGCM-1 prioritizes tool use and large context over memorization
Researchers from Zhongguancun Academy and Zhongguancun Institute of AI have developed ZGCM-1, a 7.39B parameter model that prioritizes tool use and a large context window over memorizing vast datasets. This approach all…
-
Open-source Iris search agent challenges closed-source rivals with advanced context management
AllSpark Research has launched Iris, an open-source search agent that challenges closed-source competitors. Iris utilizes a Mixture-of-Experts architecture and features a 256K context window, with models available under…
-
ByteDance's HarnessDev benchmark tests LLMs' ability to build agent code
Researchers from ByteDance Seed and other institutions have introduced HarnessDev, a new benchmark designed to evaluate an LLM's ability to create its own agent harnesses. Unlike traditional benchmarks that fix the harn…
-
OpenResearcher pipeline enables offline synthesis of AI research trajectories
Researchers have developed OpenResearcher, an open-source pipeline designed for synthesizing long-horizon research trajectories for training deep research agents. This pipeline operates offline, utilizing three explicit…
-
AllSpark Research unveils Iris, an open-weight web search agent
AllSpark Research has introduced Iris, an open-weight web search agent system designed to tackle complex, multi-hop questions that often stump current language models. Iris utilizes two models, Iris-mini and Iris-pro, p…
-
Iris search agents achieve SOTA open-source results on web benchmarks · 2 sources tracked
Researchers have developed two large-scale search agents, Iris-mini and Iris-pro, trained at 35B and 397B parameters respectively. These agents utilize a novel data pipeline and training methodology that combines superv…
-
New ICA framework improves AI agents' long-horizon information seeking
Researchers have developed a new framework called Information-Aware Credit Assignment (ICA) to improve reinforcement learning for agents that seek information over long horizons. ICA addresses the challenge of assigning…
-
Perplexity unveils local-first agent with PPLX 27B model
Perplexity has released new research detailing its "Portable Computer" agent, designed for local-first, private, and cost-effective work. This agent utilizes a post-trained PPLX 27B model, achieving high accuracy on kno…
-
New DRBENCHER benchmark tests AI agents' combined browsing and math skills
Researchers have introduced DRBENCHER, a new benchmark designed to evaluate AI agents' ability to combine web browsing with multi-step mathematical computations. Unlike previous benchmarks that assess these skills in is…
-
New research enhances AI agent memory, reasoning, and grounding
Researchers are developing advanced methods for AI agents to effectively utilize long-term memory and improve their reasoning capabilities. One approach, Query-Conditioned Reuse (QCR), focuses on how agents can adapt pa…
-
ChatGPT's Work agent mode shows mixed results on complex tasks
A recent 20-hour test of ChatGPT's agent mode, now called Work and powered by GPT-5.6, revealed mixed results. While it shows promise in tasks like data aggregation and personalized messaging, it struggles with complex …
-
New research aims to improve retrieval-augmented search agents · 2 sources tracked
Two new research papers propose methods to improve the efficiency and effectiveness of retrieval-augmented search agents. The first paper, "HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents," intro…
-
New CRISP framework trains LLM search agents to be more efficient
Researchers have introduced CRISP, a new framework designed to train more efficient deep search agents powered by large language models. Unlike previous methods that simply reduce tool usage, CRISP identifies and preser…
-
New credit assignment methods enhance AI search agent training · 3 sources tracked
Researchers have developed new methods for training long-horizon search agents, which are AI systems designed to perform complex, multi-step tasks. One approach, ABSeeker, uses Answer-Backtracked Credit Assignment (ABC)…
-
Anthropic's Claude Opus 5 excels in bio/cyber tasks, bypassing Fable 5 restrictions
Anthropic has released Claude Opus 5, positioning it as a powerful tool for computational biology and cybersecurity tasks. While the "frontier" model, Fable 5, is heavily restricted in these domains, and Mythos 5 is gat…
-
Anthropic's Claude Cookbook offers advanced agent-building recipes · 2 sources tracked
Anthropic has released a collection of resources and tutorials, dubbed the Claude Cookbook, detailing how to build and deploy advanced AI agents using its Claude models and SDKs. These resources cover a range of applica…
-
New AI agents tackle deep research and misleading web data · 4 sources tracked
Researchers have introduced AREX, a new family of recursively self-improving agents designed for deep research tasks. AREX alternates between research and self-improvement loops, using an autonomous context-update tool …
-
Agents-A1-4B model shows strong performance in long-horizon search
Agents-A1-4B, a new model developed by InternScience, demonstrates strong performance across various benchmarks, particularly in long-horizon search and agentic tasks. The model, which is based on Qwen3.7-4B, significan…
-
New STAMP method improves credit assignment for deep search agents
Researchers have introduced STAMP, a novel method for improving credit assignment in deep search agents. This approach addresses the 'reward-credit mismatch' by providing targeted credit to actions that expose supportin…
-
New framework enables AI agents to self-improve in verifiable web environments
Researchers have introduced DeepSearch-Evolve, a self-distillation framework designed to train web agents within the DeepSearch-World environment. This framework aims to overcome challenges in agent training by enabling…