Lean 4 Programming Language
PulseAugur coverage of Lean 4 Programming Language — every cluster mentioning Lean 4 Programming Language across labs, papers, and developer communities, ranked by signal.
18 day(s) with sentiment data
-
New framework offers certified training for Deep Equilibrium Networks
Researchers have developed a new framework for Deep Equilibrium Networks (DEQs) that provides certified guarantees for both inference and training. This framework uses a continuation method to achieve polynomial complex…
-
AI researchers resolve best-arm identification problem using Lean 4
Researchers have resolved conjectures regarding the instance-wise sample complexity of the best-arm identification problem. They established a lower bound related to gap entropy and introduced a single algorithm that ac…
-
AI agents form society with exploiters and whistleblowers in unsupervised study
A recent study involving 100 identical AI agents tasked with proving mathematical conjectures revealed emergent societal behaviors, including exploitation and whistleblowing. When presented with a loophole in the evalua…
-
AI assists in formalizing Dong-Yang classification of binary codes
Researchers have formalized Dong and Yang's classification of optimal (n,4) binary block codes for binary symmetric channels using the Lean 4 programming language. This machine-checked proof involved feeding the origina…
-
AI models tackle unsolved math problems with formal proofs and new discoveries
Two new research papers explore the capabilities of large language models in advanced mathematical reasoning and discovery. The first paper introduces Magenta, a system that bridges informal natural language math proble…
-
Formalizing self-dual code constructions in Lean 4
This paper introduces a formalization in Lean 4 of constructions for self-dual codes, focusing on isotropic lines. It establishes an equivalence between Chinburg and Zhang's reduction in the binary Hilbert-symbol realiz…
-
New research explores theoretical limits of neural network generalization · 4 papers
Four new research papers delve into the theoretical underpinnings of generalization in neural networks. One paper establishes a necessary and sufficient condition for provable compositional generalization, focusing on s…
-
Mistral AI releases free Leanstral-1.5 model for formal proof engineering
Mistral AI has released Leanstral-1.5, a 119 billion parameter model optimized for automated theorem proving and the Lean 4 programming language. This model is available for free and aims to assist users in formally pro…
-
New StochBench benchmark tests LLMs on stochastic processes in Lean · 2 sources tracked
Researchers have introduced StochBench, a new benchmark designed to evaluate large language models on stochastic processes in the Lean 4 programming language. This benchmark features 450 graduate-level problems, address…
-
Fermat's Last Theorem formalized in Lean 4; Trump admin fights ABC lawsuit
A discussion on Mastodon highlights two distinct news items: the formalization of Fermat's Last Theorem using the Lean 4 programming language, and the Trump administration's legal battle with ABC, with concerns that Dis…
-
Anthropic's Claude AI formalizes Fermat's Last Theorem proof · 8 sources tracked
Anthropic's AI model, Claude, has successfully formalized a complete proof of Fermat's Last Theorem using the Lean proof assistant. This significant achievement, completed over 11 days, involved generating millions of l…
-
Anthropic's Claude AI achieves first computer-checked proof of Fermat's Last Theorem · 4 sources tracked
Anthropic has announced the first complete, computer-checked formalization of Fermat's Last Theorem using the Lean 4 programming language. An internal research model, built on Claude, autonomously worked for 11 days to …
-
AI Astra to explore mathematical discovery with formal verification
A scientist is proposing an experimental framework called Astra to explore AI's potential in mathematical discovery, moving beyond simple pattern retrieval. The core idea is to create a loop where an AI, Astra, explores…
-
AI system AutoGraphForge automates graph theory conjecture discovery and proof
Researchers have developed AutoGraphForge, a computational pipeline designed to automate the discovery and proving of graph theory conjectures. The system generates conjectures using a counterexample-guided approach, fi…
-
New framework TopoAlign uses code to boost LLM math reasoning
Researchers have introduced TopoAlign, a novel framework designed to enhance the mathematical reasoning capabilities of large language models (LLMs) by leveraging vast code repositories. This approach addresses the scar…
-
New SHADOWBENCH benchmark improves evaluation of AI-generated math code
Researchers have introduced SHADOWBENCH, a new benchmark designed to more reliably evaluate the semantic alignment of autoformalized mathematical statements. This benchmark utilizes a novel metric called SA-Pass, which …
-
AI-generated math proofs create verification abundance, adjudication scarcity
A recent arXiv paper explores the implications of AI models generating mathematically verifiable proofs, highlighting a shift in the verification process. While AI can now produce machine-checkable proofs, the paper arg…
-
New MCTS framework for AI theorem proving highlights need for proof auditing
Researchers have developed a novel three-role Monte Carlo Tree Search (MCTS) framework for formal theorem proving using large language models. This approach treats the Lean 4 compiler as a reward oracle, using its outpu…
-
New platform Prove2Me enables AI-assisted collaborative math formalization · 4 sources tracked
Researchers have introduced Prove2Me, an open platform designed to facilitate large-scale, collaborative formalization of mathematics. This platform leverages AI coding agents to assist human users in writing formal pro…
-
New MathAdv Benchmark Evaluates Theorem Prover Reasoning Capabilities
A new benchmark called MathAdv has been developed to evaluate the mathematical reasoning capabilities of theorem provers. This benchmark spans 13 mathematical domains and includes auxiliary tasks such as multiple-choice…