Researchers have introduced Iris-mini and Iris-pro, two search agents trained at 35B-A3B and 397B-A17B scales, respectively. These agents are developed using a novel data pipeline that constructs multi-hop chains over web corpus entity graphs and generates questions that are challenging for reference models. The training process involves alternating supervised fine-tuning (SFT) with reinforcement learning (RL) against live search, a method termed SFT-RL climbing. When evaluated with context management, Iris-pro achieved strong results on benchmarks like BrowseComp and DeepSearchQA, outperforming other open-source search agents in its parameter range. AI
IMPACT These models advance the capabilities of open-source search agents, potentially improving web navigation and information retrieval.
RANK_REASON The item is a research paper detailing new AI models and training methodologies. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- BrowseComp+
- BrowseComp-ZH
- CatalyzeX
- CORE Recommender
- DagsHub
- DeepSearchQA
- Gotit.pub
- Hugging Face
- Influence Flower
- Iris Minich
- React
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →