PulseAugur
实时 10:28:21
English(EN) Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

AI模型在新的HorizonMath基准上发现新的数学解决方案

研究人员推出了HorizonMath,这是一个旨在评估AI解决复杂、未解决数学问题能力的新基准。该基准包含八个领域的113个问题,重点关注“生成器-验证器差距”,即发现困难但验证简单的领域。初步测试表明,大多数当前模型得分低于10%,但GPT-5.4 Pro和GPT-5.6 Sol各自成功发现了三个研究问题的新颖解决方案,为数学文献做出了贡献。 AI

影响 为评估AI的数学推理和发现能力设定了新标准,有可能加速AI驱动的科学研究的进展。

排序理由 该集群是关于一篇介绍基准并报告研究结果的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AI模型在新的HorizonMath基准上发现新的数学解决方案

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群是关于一篇介绍基准并报告研究结果的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles London, Kalyan Ramakrishnan, Jakob Foerster, Cheng Zhang, Flaviu Cipcigan, Philip Torr, Alessandro Abate ·

    使用自动验证衡量在数学发现推理方面的进展

    arXiv:2603.15617v2 Announce Type: replace Abstract: Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated…