PulseAugur
EN
LIVE 11:21:57

Language models prove 30% of math conjectures in new benchmark

Researchers have developed OEIS Open, a new benchmark designed to evaluate how many mathematical conjectures language models can prove. This benchmark, based on 492 open conjectures from the On-Line Encyclopedia of Integer Sequences formalized in Lean, allows any generic language model to be tested. Initial results show that LMs can resolve a significant portion of these conjectures autonomously, with one model scoring 44% on a subset of the benchmark using a budget of $200 per attempt. AI

IMPACT Demonstrates potential for LLMs to autonomously discover mathematical proofs, potentially accelerating research in formal mathematics.

RANK_REASON The cluster describes a new benchmark for evaluating language models on mathematical conjectures, based on an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Language models prove 30% of math conjectures in new benchmark

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tom Adamczewski ·

    OEIS Open: How many conjectures can language models turn into theorems?

    arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source …