DeepSeek-Prover-V2-7B
PulseAugur coverage of DeepSeek-Prover-V2-7B — every cluster mentioning DeepSeek-Prover-V2-7B across labs, papers, and developer communities, ranked by signal.
-
Flaws Found in Lean Theorem Proving Benchmarks and RL Model Inference
Researchers have identified significant flaws in the formal benchmarking of Lean theorem-proving datasets, uncovering thousands of issues including counterexamples and vacuous theorems. A separate study on RL-trained Le…
-
New LLM Frameworks and Benchmarks Advance Formal Mathematical Reasoning
Researchers are developing new methods and benchmarks to improve the formal mathematical reasoning capabilities of large language models (LLMs). One approach, Diffusion-Proof, utilizes diffusion LLMs (dLLMs) for theorem…
-
FormalRewardBench benchmark evaluates LLM reward models for theorem proving
Researchers have introduced FormalRewardBench, a new benchmark designed to evaluate reward models used in formal theorem proving. This benchmark addresses the challenge of sparse credit assignment in reinforcement learn…