PulseAugur
实时 10:29:33
English(EN) OEIS Open: How many conjectures can language models turn into theorems?

语言模型在新基准测试中证明了30%的数学猜想

研究人员开发了OEIS Open,这是一个旨在评估语言模型能证明多少数学猜想的新基准。该基准基于来自整数序列在线百科(On-Line Encyclopedia of Integer Sequences)的492个已形式化为Lean的开放猜想,允许测试任何通用语言模型。初步结果表明,语言模型能够自主解决其中很大一部分猜想,一个模型在使用每次200美元的预算在一个子集上取得了44%的得分。 AI

影响 展示了大型语言模型自主发现数学证明的潜力,可能加速形式数学的研究。

排序理由 该集群基于一篇学术论文,描述了一个用于评估语言模型在数学猜想方面的能力的新基准测试。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

语言模型在新基准测试中证明了30%的数学猜想

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tom Adamczewski ·

    OEIS公开:语言模型能将多少猜想转化为定理?

    arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source …