Researchers have developed new tools and benchmarks for formal theorem proving, an area increasingly relevant to AI. One paper details an interactive sequent prover for Event-B, encoded in Prolog, which offers advantages in teaching and proof analysis. Another paper presents extensions to ProB, a Prolog-based model checker, for animating and visualizing Prolog transition systems, with applications in game strategy evaluation and teaching. A third contribution introduces ITPEval, the first benchmark for translating formal proofs across different interactive theorem provers (ITPs), revealing that library mismatches are a significant bottleneck for LLM-based translation. Finally, a new agent called AoA operates on abstract syntax trees rather than concrete syntax, significantly reducing API costs and improving efficiency for LLM-based proof agents, while also enabling the use of newer proof languages. AI
IMPACT These advancements in formal proof systems and benchmarks could accelerate AI's capabilities in program verification and formalized mathematics.
RANK_REASON Multiple research papers published on arXiv detailing new tools, benchmarks, and methods in formal theorem proving and AI.
- Isabelle Agent
- Joshua Jun Leang Ong
- miniF2F
- Minilang
- NTP4VC-Pearl
- Agent over AST (AoA)
- Beq
- HOL Light
- Isabelle
- ITPEval
- Lean 4 Programming Language
- Rocq prover
- arXiv
- Prolog
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →