Researchers have developed OEIS Open, a new benchmark designed to evaluate how many mathematical conjectures language models can prove. The benchmark, based on 492 open conjectures from the On-Line Encyclopedia of Integer Sequences formalized in Lean, allows generic language models to attempt proofs autonomously. Initial results show that LMs can resolve a significant portion of these conjectures at a modest cost, with one model scoring 44% on a subset of the benchmark using a $200 budget per attempt. AI
IMPACT Demonstrates potential for LMs to autonomously resolve open research conjectures, potentially accelerating mathematical discovery.
RANK_REASON The cluster describes a new benchmark and research paper evaluating language models on mathematical conjectures.
Read on Hugging Face Daily Papers →
- arXiv
- Language Models
- Lean
- OEIS Open
- On-Line Encyclopedia of Integer Sequences
- Thomas Adamczewski
- Tsoukalas et al.
- OEIS Open Lite
- Tsoukalas
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →