Researchers have developed ProofJudge, an LLM-based system designed to evaluate the quality of formal proofs written in the Lean 4 programming language, specifically within the Mathlib library. This agentic system assesses proofs across five criteria beyond mere correctness, including library leverage, automation fit, structural clarity, statement quality, and adherence to Mathlib conventions. ProofJudge was evaluated on a dataset of 218 Mathlib pull requests, demonstrating an ability to align with human reviewer preferences significantly above chance, with some open-weight models achieving high accuracy at a reduced cost. The project aims to foster further research by releasing the judge harness, evaluation dataset, and traces as open-source artifacts. AI
IMPACT This tool could improve the efficiency and consistency of code review for formal mathematical proofs, potentially accelerating development in theorem proving communities.
RANK_REASON The cluster describes a new research paper detailing an LLM-based evaluation system for formal proofs.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →