DeepSeek-Prover-V2-7B
PulseAugur coverage of DeepSeek-Prover-V2-7B — every cluster mentioning DeepSeek-Prover-V2-7B across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New SCOPE system uses LLM for certified theorem proving
A new research paper introduces SCOPE, a system designed to improve certified theorem proving in proof assistants like Lean. SCOPE divides labor between a language model that plans proof steps and a symbolic engine that…
-
New MCTS framework for AI theorem proving highlights need for proof auditing
Researchers have developed a novel three-role Monte Carlo Tree Search (MCTS) framework for formal theorem proving using large language models. This approach treats the Lean 4 compiler as a reward oracle, using its outpu…
-
Flaws Found in Lean Theorem Proving Benchmarks and RL Model Inference
Researchers have identified significant flaws in the formal benchmarking of Lean theorem-proving datasets, uncovering thousands of issues including counterexamples and vacuous theorems. A separate study on RL-trained Le…
-
New LLM Frameworks and Benchmarks Advance Formal Mathematical Reasoning
Researchers are developing new methods and benchmarks to improve the formal mathematical reasoning capabilities of large language models (LLMs). One approach, Diffusion-Proof, utilizes diffusion LLMs (dLLMs) for theorem…
-
FormalRewardBench benchmark evaluates LLM reward models for theorem proving
Researchers have introduced FormalRewardBench, a new benchmark designed to evaluate reward models used in formal theorem proving. This benchmark addresses the challenge of sparse credit assignment in reinforcement learn…