Researchers have introduced AREX, a new family of recursively self-improving agents designed for deep research tasks. AREX alternates between research and self-improvement loops, using an autonomous context-update tool to manage growing interaction history. This approach allows AREX to outperform comparable-scale baselines on benchmarks like BrowseComp and Humanity's Last Exam. Concurrently, a separate study introduces DRNOISE, a benchmark designed to evaluate deep research agents' ability to handle misleading information on the open web, highlighting significant accuracy drops when such documents are present. AI
IMPACT These developments highlight advancements in AI agent capabilities for complex research and their robustness against misleading information.
RANK_REASON The cluster contains two research papers introducing new AI agent architectures and benchmarks.
Read on arXiv cs.IR (Information Retrieval) →
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DRNOISE
- Gotit.pub
- Hugging Face
- IArxiv
- ScienceCast
- AREX
- Beijing Academy of Artificial Intelligence
- BrowseComp
- Deep Research
- DeepSearchQA
- Humanity's Last Exam
- Recursively Self-Improving
- WideSearch
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →