PulseAugur
实时 05:14:54
English(EN) PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

AI代理解决复杂数学问题,树立新的研究基准 · 追踪8个来源

研究人员正在开发能够解决复杂数学问题的先进AI代理,拓展自动化推理的边界。ProofCouncil和OpenProver等系统在解决开放性数学问题和生成形式化证明方面展现出显著能力,其中ProofCouncil在涉及10个现实世界问题的挑战中取得了显著成功。IMProofBench和MIRA-Math等新基准支持了这些努力,这些基准旨在严格评估LLM在研究级数学任务上的表现以及它们请求必要信息的能力。 AI

影响 AI在数学领域的进步可能加速科学发现和定理证明。

排序理由 多篇研究论文介绍了用于数学推理的新AI系统和基准。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 21 个来源。 我们如何撰写摘要 →

AI代理解决复杂数学问题,树立新的研究基准 · 追踪8个来源

报道来源 [21]

  1. arXiv cs.CL TIER_1 English(EN) · Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen ·

    AdvancedMathBench:用于高级数学证明生成和验证的基准套件

    arXiv:2607.11849v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in bo…

  2. arXiv cs.CL TIER_1 English(EN) · Burak S. Akbudak, Zeynel A. Ulu\c{s}an, Can S. Erer, G\"ozde G\"ul \c{S}ahin ·

    TreeThink:用于LLM数学推理的模块化树搜索库

    arXiv:2607.11258v1 Announce Type: new Abstract: Tree search algorithms enable systematic exploration of the proof space in neural theorem proving. Existing LLM tree search libraries primarily target natural language reasoning and do not provide native integration with formal veri…

  3. arXiv cs.CL TIER_1 English(EN) · Kai Chen ·

    AdvancedMathBench:用于高级数学证明生成和验证的基准套件

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvancedMathBench:用于高级数学证明生成和验证的基准套件

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvancedMathBench:用于高级数学证明生成和验证的基准套件

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  6. arXiv cs.CL TIER_1 English(EN) · Gözde Gül Şahin ·

    TreeThink:用于LLM数学推理的模块化树搜索库

    Tree search algorithms enable systematic exploration of the proof space in neural theorem proving. Existing LLM tree search libraries primarily target natural language reasoning and do not provide native integration with formal verifiers, while theorem proving systems often rely …

  7. arXiv cs.AI TIER_1 English(EN) · Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely B\'erczi, Uri Kreitner, Liam Price, David Holmes ·

    ProofCouncil:一个用于解决开放性数学问题的LLM智能体

    arXiv:2607.09474v1 Announce Type: new Abstract: Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this e…

  8. arXiv cs.AI TIER_1 English(EN) · Mat\v{e}j Kripner, Milan Straka ·

    OpenProver:使用 Lean 4 进行基于代理的交互式定理证明

    arXiv:2607.09217v1 Announce Type: new Abstract: In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by r…

  9. arXiv cs.AI TIER_1 English(EN) · David Holmes ·

    ProofCouncil:一个用于解决开放性数学问题的LLM智能体

    Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical ag…

  10. arXiv cs.AI TIER_1 English(EN) · Milan Straka ·

    OpenProver:使用 Lean 4 进行基于代理和交互式的定理证明

    In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia. A Pl…

  11. arXiv cs.AI TIER_1 English(EN) · Eric Jiang, Xiao Liang, Yikai Zhang, Yingjia Wan, Mengting Li, Haikang Deng, Alexander K. Taylor, Justin Baker, Rushil Raghavan, Junyi Zhang, Ying Nian Wu, Andrea L. Bertozzi, Kai-Wei Chang, Raghu Meka, Matthew Sottile, Nanyun Peng, Amit Sahai, Terence T… ·

    从求解器到研究:大型语言模型驱动的数理研究前沿

    arXiv:2607.07779v1 Announce Type: cross Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interacti…

  12. arXiv cs.CL TIER_1 English(EN) · Johannes Schmitt, Gergely B\'erczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, Jo\~ao Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontana… ·

    IMProofBench:对AI进行研究级数学证明生成基准测试

    arXiv:2509.26076v2 Announce Type: replace Abstract: As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing bench…

  13. arXiv cs.AI TIER_1 English(EN) · Pavel Snopov, German Magai ·

    评估SageMath增强的LLM代理在计算和实验数学中的应用

    arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup t…

  14. arXiv cs.AI TIER_1 English(EN) · Charbel Al Bateh, Samer Saab Jr ·

    MIRA-Math:最小信息请求和数学推理的基准测试

    arXiv:2607.07391v1 Announce Type: new Abstract: Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a…

  15. arXiv cs.CL TIER_1 English(EN) · Wei Wang ·

    从求解器到研究:大型语言模型驱动的数学形式化研究前沿

    Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, curre…

  16. arXiv cs.AI TIER_1 English(EN) · Samer Saab ·

    MIRA-Math:最小信息请求和数学推理的基准测试

    Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathema…

  17. arXiv cs.AI TIER_1 English(EN) · Daryna Dementieva, Nikolay Babakov, Kathy H\"ammerl, Ilseyar Alimova, Jind\v{r}ich Libovick\'y, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan \"Ozer, Nikola Selic, Subhankar Swain, Tsedeniya Kin… ·

    PluraMath: 将数学推理评估扩展到高资源语言之外

    arXiv:2607.05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating b…

  18. arXiv cs.AI TIER_1 English(EN) · Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima ·

    大型语言模型中的数学推理:基准、架构、评估和开放性挑战

    arXiv:2605.19723v2 Announce Type: replace-cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large Language Models (LLMs) improve their reas…

  19. arXiv cs.AI TIER_1 English(EN) · Alexander Fraser ·

    PluraMath: 将数学推理评估扩展到高资源语言之外

    Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. Th…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    PluraMath: 将数学推理评估扩展到高资源语言之外

    PluraMath extends the PolyMath dataset to 18 underrepresented languages, revealing persistent gaps in multilingual mathematical reasoning performance between high-resource and low-resource languages.

  21. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    AdvancedMathBench:LLM高级数学推理的新基准

    <h2> What Changed </h2> <p>Large language models (LLMs) have demonstrated proficiency in high-school and olympiad-style mathematics. However, their performance in advanced mathematics has remained less understood due to limitations in existing benchmarks. These prior benchmarks o…