Researchers have introduced Dr. Zero, a novel framework designed to enable large language model (LLM) search agents to self-evolve without relying on curated training data. This approach utilizes an external search engine as the agent's knowledge environment and incorporates a self-evolution feedback loop where a proposer agent generates diverse questions to train a solver agent. To improve training efficiency, the framework employs hop-grouped relative policy optimization (HRPO), which clusters similar questions to reduce computational overhead. Experiments show that Dr. Zero can match or exceed the performance of supervised search agents on question-answering benchmarks, demonstrating the potential for strong agentic search and reasoning to emerge solely through self-evolution. AI
IMPACT This research could significantly reduce the need for large, curated datasets in training AI agents, potentially accelerating development and deployment.
RANK_REASON Academic paper detailing a new AI framework and method. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →