PulseAugur
EN
LIVE 07:48:25

AI agents tackle complex math problems, setting new research benchmarks · 8 sources tracked

Researchers are developing advanced AI agents capable of tackling complex mathematical problems, pushing the boundaries of automated reasoning. Systems like ProofCouncil and OpenProver are demonstrating significant capabilities in solving open mathematical problems and generating formal proofs, with ProofCouncil achieving notable success in a challenge involving 10 real-world problems. These efforts are supported by new benchmarks such as IMProofBench and MIRA-Math, designed to rigorously evaluate LLMs on research-level mathematical tasks and their ability to request necessary information. AI

IMPACT Advances in AI for mathematics could accelerate scientific discovery and theorem proving.

RANK_REASON Multiple research papers introducing new AI systems and benchmarks for mathematical reasoning.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 21 sources. How we write summaries →

AI agents tackle complex math problems, setting new research benchmarks · 8 sources tracked

COVERAGE [21]

  1. arXiv cs.CL TIER_1 English(EN) · Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen ·

    AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

    arXiv:2607.11849v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in bo…

  2. arXiv cs.CL TIER_1 English(EN) · Burak S. Akbudak, Zeynel A. Ulu\c{s}an, Can S. Erer, G\"ozde G\"ul \c{S}ahin ·

    TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs

    arXiv:2607.11258v1 Announce Type: new Abstract: Tree search algorithms enable systematic exploration of the proof space in neural theorem proving. Existing LLM tree search libraries primarily target natural language reasoning and do not provide native integration with formal veri…

  3. arXiv cs.CL TIER_1 English(EN) · Kai Chen ·

    AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

    Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provid…

  6. arXiv cs.CL TIER_1 English(EN) · Gözde Gül Şahin ·

    TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs

    Tree search algorithms enable systematic exploration of the proof space in neural theorem proving. Existing LLM tree search libraries primarily target natural language reasoning and do not provide native integration with formal verifiers, while theorem proving systems often rely …

  7. arXiv cs.AI TIER_1 English(EN) · Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely B\'erczi, Uri Kreitner, Liam Price, David Holmes ·

    ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

    arXiv:2607.09474v1 Announce Type: new Abstract: Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this e…

  8. arXiv cs.AI TIER_1 English(EN) · Mat\v{e}j Kripner, Milan Straka ·

    OpenProver: Agentic and Interactive Theorem Proving with Lean 4

    arXiv:2607.09217v1 Announce Type: new Abstract: In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by r…

  9. arXiv cs.AI TIER_1 English(EN) · David Holmes ·

    ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

    Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical ag…

  10. arXiv cs.AI TIER_1 English(EN) · Milan Straka ·

    OpenProver: Agentic and Interactive Theorem Proving with Lean 4

    In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia. A Pl…

  11. arXiv cs.AI TIER_1 English(EN) · Eric Jiang, Xiao Liang, Yikai Zhang, Yingjia Wan, Mengting Li, Haikang Deng, Alexander K. Taylor, Justin Baker, Rushil Raghavan, Junyi Zhang, Ying Nian Wu, Andrea L. Bertozzi, Kai-Wei Chang, Raghu Meka, Matthew Sottile, Nanyun Peng, Amit Sahai, Terence T… ·

    From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

    arXiv:2607.07779v1 Announce Type: cross Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interacti…

  12. arXiv cs.CL TIER_1 English(EN) · Johannes Schmitt, Gergely B\'erczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, Jo\~ao Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontana… ·

    IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

    arXiv:2509.26076v2 Announce Type: replace Abstract: As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing bench…

  13. arXiv cs.AI TIER_1 English(EN) · Pavel Snopov, German Magai ·

    Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

    arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup t…

  14. arXiv cs.AI TIER_1 English(EN) · Charbel Al Bateh, Samer Saab Jr ·

    MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

    arXiv:2607.07391v1 Announce Type: new Abstract: Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a…

  15. arXiv cs.CL TIER_1 English(EN) · Wei Wang ·

    From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

    Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, curre…

  16. arXiv cs.AI TIER_1 English(EN) · Samer Saab ·

    MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

    Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathema…

  17. arXiv cs.AI TIER_1 English(EN) · Daryna Dementieva, Nikolay Babakov, Kathy H\"ammerl, Ilseyar Alimova, Jind\v{r}ich Libovick\'y, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan \"Ozer, Nikola Selic, Subhankar Swain, Tsedeniya Kin… ·

    PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    arXiv:2607.05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating b…

  18. arXiv cs.AI TIER_1 English(EN) · Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima ·

    Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

    arXiv:2605.19723v2 Announce Type: replace-cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large Language Models (LLMs) improve their reas…

  19. arXiv cs.AI TIER_1 English(EN) · Alexander Fraser ·

    PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. Th…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    PluraMath extends the PolyMath dataset to 18 underrepresented languages, revealing persistent gaps in multilingual mathematical reasoning performance between high-resource and low-resource languages.

  21. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

    <h2> What Changed </h2> <p>Large language models (LLMs) have demonstrated proficiency in high-school and olympiad-style mathematics. However, their performance in advanced mathematics has remained less understood due to limitations in existing benchmarks. These prior benchmarks o…