PulseAugur
EN
LIVE 08:20:42

AI agent Intern-S1-MO tackles Olympiad-level math problems

Researchers have developed Intern-S1-MO, a novel long-horizon reasoning agent designed to tackle complex mathematical problems at the Olympiad level. This agent employs a multi-round, hierarchical reasoning approach using a system of interconnected Large Reasoning Models (LRMs) for reasoning, summarization, and verification. By maintaining a compact memory of lemmas, Intern-S1-MO can explore extensive reasoning spaces across multiple stages, overcoming the context length limitations of standard LRMs. The system is trained using a new Reinforcement Learning framework, OREAL-H, which simultaneously enhances the LRM's reasoning capabilities and the agent's overall performance. Experiments demonstrate that Intern-S1-MO achieves significant results, matching silver medalists on IMO2025 non-geometry problems and reaching gold medal level in the CMO2025 competition. AI

IMPACT This research demonstrates a novel approach to long-horizon reasoning in AI, potentially advancing capabilities in complex problem-solving domains.

RANK_REASON The item is a research paper detailing a new AI agent and training framework for mathematical problem-solving. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AI agent Intern-S1-MO tackles Olympiad-level math problems

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yuzhe Gu, Songyang Gao, Zijian Wu, Lingkai Kong, Wenwei Zhang, Zhongrui Cai, Fan Zheng, Tianyou Ma, Junhao Shen, Haiteng Zhao, Duanyang Zhang, Huilun Zhang, Kuikun Liu, Chengqi Lyu, Yanhui Duan, Chiyu Chen, Ningsheng Ma, Jianfei Gao, Han Lyu, Dahua Lin, … ·

    Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

    arXiv:2512.10739v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning with Verifiable Rewards (RLVR), capable of solving AIME-level problems. However…