miniF2F-test
PulseAugur coverage of miniF2F-test — every cluster mentioning miniF2F-test across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Flaws Found in Lean Theorem Proving Benchmarks and RL Model Inference
Researchers have identified significant flaws in the formal benchmarking of Lean theorem-proving datasets, uncovering thousands of issues including counterexamples and vacuous theorems. A separate study on RL-trained Le…
-
New matrix refines LLM autoformalization error analysis
Researchers have introduced a "signal-coverage matrix" to better evaluate the performance of Large Language Models (LLMs) in autoformalization tasks. This matrix stratifies errors into type-correctness and semantic-equi…
-
New LLM Frameworks and Benchmarks Advance Formal Mathematical Reasoning
Researchers are developing new methods and benchmarks to improve the formal mathematical reasoning capabilities of large language models (LLMs). One approach, Diffusion-Proof, utilizes diffusion LLMs (dLLMs) for theorem…
-
Pythagoras-Prover achieves state-of-the-art in efficient formal proving
Researchers have introduced Pythagoras-Prover, a new family of theorem provers designed for efficiency in formal reasoning tasks. These models utilize curriculum training and augmented formalization techniques to overco…
-
AI frameworks boost formal theorem proving with new techniques
Researchers have developed new frameworks to enhance formal theorem proving capabilities using large language models. Goedel-Architect utilizes a blueprint generation and refinement strategy, achieving state-of-the-art …
-
Knowledge graphs boost LLMs for automated theorem proving
Researchers have developed KG-Prover, a new framework that enhances large language models for automated theorem proving by integrating knowledge graphs mined from mathematical texts. This approach helps LLMs identify ke…