PulseAugur
实时 07:11:46
English(EN) ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

新的ScienceArena基准测试大型语言模型在奥林匹克级别科学问题上的表现

研究人员开发了ScienceArena,这是一个旨在评估大型语言模型(LLMs)科学推理能力的新基准。该基准借鉴了最近的物理、化学和生物学竞赛,包括国际物理奥林匹克竞赛和国际化学奥林匹克竞赛。ScienceArena采用过程信用评分系统,并由专家和奥林匹克奖牌获得者进行数字化和验证,以确保准确性。对十四个LLMs的初步评估表明,虽然顶级模型在某些考试中取得了奖牌级别的分数,但在化学和在长串解题过程中保持一致性方面仍然存在挑战。 AI

影响 该基准测试有望推动LLMs在科学推理和解决问题能力方面的改进。

排序理由 该集群包含一篇介绍用于评估LLMs的新基准的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的ScienceArena基准测试大型语言模型在奥林匹克级别科学问题上的表现

本文如何被排名

Signal score
24 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇介绍用于评估LLMs的新基准的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yuhan Wu, Tong Yang, Lin Sun, Xiangzheng Zhang ·

    ScienceArena:在最新科学奥林匹克竞赛中对大型语言模型进行基准测试

    arXiv:2608.30517v1 Announce Type: new Abstract: Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, ch…