Researchers have developed Intern-S1-MO, a novel long-horizon reasoning agent designed to tackle complex mathematical problems at the Olympiad level. This agent employs a multi-round, hierarchical reasoning approach using a system of interconnected Large Reasoning Models (LRMs) for reasoning, summarization, and verification. By maintaining a compact memory of lemmas, Intern-S1-MO can explore extensive reasoning spaces across multiple stages, overcoming the context length limitations of standard LRMs. The system is trained using a new Reinforcement Learning framework, OREAL-H, which simultaneously enhances the LRM's reasoning capabilities and the agent's overall performance. Experiments demonstrate that Intern-S1-MO achieves significant results, matching silver medalists on IMO2025 non-geometry problems and reaching gold medal level in the CMO2025 competition. AI
IMPACT This research demonstrates a novel approach to long-horizon reasoning in AI, potentially advancing capabilities in complex problem-solving domains.
RANK_REASON The item is a research paper detailing a new AI agent and training framework for mathematical problem-solving. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →