PulseAugur
中
实时 02:44:08
English(EN) Distilling LLM Feedback for Lean Theorem Proving

新基准评估大型语言模型的数学推理和证明验证能力

研究人员引入了新的基准和评估方法来评估大型语言模型的数学推理能力。ComBench 专注于奥林匹克级别的组合数学,区分证明推理和构造性实现,并发现即使是顶级模型也难以胜任这些复杂任务。另一种方法 TheoremBench 使用 Lean4 语言在形式数学中评估大型语言模型的定理证明能力,强调需要超越竞赛式问题来评估模型在更长、依赖性更强的数学发展中的表现。此外,一种用于严格逐级验证研究级证明的方法旨在通过仔细检查每个推理步骤来解决大型语言模型不可靠的问题。 AI

影响 这些基准和验证方法将推动大型语言模型在数学推理和形式证明能力方面的进步。

排序理由 多篇研究论文介绍了用于评估大型语言模型在数学推理方面能力的新基准和评估方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

新基准评估大型语言模型的数学推理和证明验证能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了用于评估大型语言模型在数学推理方面能力的新基准和评估方法。
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
132 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+6 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [13]

  1. arXiv cs.AI TIER_1 English(EN) · Jiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng… ·

    MaxProof:使用生成式验证器强化学习和群体级测试时缩放来扩展数学证明

    arXiv:2606.13473v1 Announce Type: cross Abstract: We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and …

  2. arXiv cs.AI TIER_1 English(EN) · Joshua Ong Jun Leang, Zheng Zhao, Mihaela C\u{a}t\u{a}lina Stoian, Qiyuan Xu, Haonan Li, Wenda Li, Shay B. Cohen, Eleonora Giunchiglia ·

    Pythagoras-Prover:通过增强的Lean形式化推进高效形式证明

    arXiv:2606.12594v1 Announce Type: new Abstract: Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven in part by scarce verified proof data and the long reasoning traces of formal proof search, making both supervised f…

  3. arXiv cs.AI TIER_1 English(EN) · Yu Cheng ·

    MaxProof:利用生成-验证器强化学习和群体级测试时缩放来扩展数学证明

    We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defen…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    MaxProof:使用生成式验证器强化学习和群体级测试时缩放来扩展数学证明

    MaxProof is a test-time scaling framework that enhances mathematical proof generation by combining multiple proof-oriented capabilities and using population-level search with tournament selection to achieve competitive performance on high-level mathematical competitions.

  5. arXiv cs.AI TIER_1 English(EN) · Yifeng Sun ·

    通过严格的逐级验证评估研究级数学证明

    arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from "context poisoning," in which superficially plausible statements mask subtle logical flaws, le…

  6. arXiv cs.AI TIER_1 English(EN) · Shunkai Zhang, Haoran Zhang, Yun Luo, Qianjia Cheng, Haodi Lei, Yizhuo Li, Runzhe Zhan, Zhilin Wang, Bangjie Xu, Yucheng Su, Xinmiao Han, Xiaoye Qu, Dongrui Liu, Zhouchen Lin, Yu Qiao, Ning Ding, Yafu Li, Yu Cheng ·

    ComBench:一个用于奥林匹克级别组合学中严谨证明推理和构造性实现的基准测试

    arXiv:2606.10479v1 Announce Type: new Abstract: Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight. Recent evidence suggests that even today's strongest frontier model…

  7. arXiv cs.AI TIER_1 English(EN) · Yifeng Sun ·

    通过严格的逐级验证评估研究级数学证明

    Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from "context poisoning," in which superficially plausible statements mask subtle logical flaws, leading to hallucination or over-skepticism. To ad…

  8. arXiv cs.AI TIER_1 English(EN) · George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Anja Surina, Moritz Firsching, Gergely B\'erczi, Francisco J. R. Ruiz, Arun Suggala, Adam Zsolt Wagner, Eric Wieser, Lei Yu, Aja Huang, Mikl\'os Z. Horv\'ath, Andrew Ferraiuolo, Henryk Michalewski, Ed… ·

    利用人工智能驱动的自动定理证明搜索推进数学研究

    arXiv:2605.22763v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research. A mitigation is using LLMs to generate formal proofs in languages like Lean. We per…

  9. arXiv cs.AI TIER_1 English(EN) · QuocViet Pham, Elvir Karimov, Andrey Galichin, Ivan Oseledets ·

    TheoremBench:在形式数学中评估LLM的定理证明能力

    arXiv:2606.09450v1 Announce Type: new Abstract: LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to capture how models behave on longer, more dependency-…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    ComBench:用于奥赛级别组合数学中严谨证明推理和构造性实现的基准测试

    A new benchmark called ComBench is introduced to evaluate large language models' combinatorial reasoning abilities through Olympiad-level problems that test both proof construction and explicit mathematical constructions.

  11. arXiv cs.AI TIER_1 English(EN) · Ivan Oseledets ·

    TheoremBench:在形式数学中评估LLM的定理证明能力

    LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to capture how models behave on longer, more dependency-rich mathematical developments. We introduce The…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    TheoremBench:在形式数学中评估LLM的定理证明能力

    LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to capture how models behave on longer, more dependency-rich mathematical developments. We introduce The…

  13. Hugging Face Daily Papers TIER_1 English(EN) ·

    提炼大型语言模型反馈以实现精简定理证明

    Feedback Distillation improves post-training of reasoning models by using self-distillation with token-level supervision and privileged feedback from language models, offering better diversity and complementary benefits when combined with GRPO.