PulseAugur
实时 06:30:29
English(EN) DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

新基准测试大型语言模型在全新物理世界中的科学发现能力

研究人员开发了DiscoverPhysics,这是一个旨在测试大型语言模型(LLM)科学推理能力的新基准。这个交互式基准挑战LLM代理在具有新颖物理学的模拟世界中发现运动定律,要求它们设计实验、观察数据,并提出自然语言解释和推断定律的Python实现。对十一个前沿模型的评估显示,即使是最强的代理也只能通过一半的世界,尤其是在揭示潜在结构方面遇到困难。该基准还突显了开源模型和商业模型之间存在的显著差距,后者在实验设计和数据提取能力方面表现更优。 AI

影响 该基准有望推动具有更强大科学推理和假设生成能力的大型语言模型的发展。

排序理由 该集群描述了一篇介绍用于评估大型语言模型能力的新颖基准的学术论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准测试大型语言模型在全新物理世界中的科学发现能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇介绍用于评估大型语言模型能力的新颖基准的学术论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
99 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Matt L. Wiemann, Lindsay M. Smith, Peter Melchior, Siddharth Mishra-Sharma, Andrew Gordon Wilson, Pavel Izmailov, Carolina Cuesta-L\'azaro ·

    DiscoverPhysics:对开箱即用的科学思维进行大语言模型基准测试

    arXiv:2605.26087v1 Announce Type: cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a …

  2. arXiv stat.ML TIER_1 English(EN) · Carolina Cuesta-Lázaro ·

    DiscoverPhysics:为开箱即用的科学思维进行大语言模型基准测试

    Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simu…