Researchers have developed OEIS Open, a new benchmark designed to evaluate how many mathematical conjectures language models can prove. This benchmark, based on 492 open conjectures from the On-Line Encyclopedia of Integer Sequences formalized in Lean, allows any generic language model to be tested. Initial results show that LMs can resolve a significant portion of these conjectures autonomously, with one model scoring 44% on a subset of the benchmark using a budget of $200 per attempt. AI
IMPACT Demonstrates potential for LLMs to autonomously discover mathematical proofs, potentially accelerating research in formal mathematics.
RANK_REASON The cluster describes a new benchmark for evaluating language models on mathematical conjectures, based on an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Language Models
- Lean
- OEIS Open
- On-Line Encyclopedia of Integer Sequences
- Thomas Adamczewski
- Tsoukalas et al.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →