PulseAugur
中
实时 10:42:05
English(EN) SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning

新的SpeedrunBench用视频游戏速通挑战大型语言模型代理

研究人员推出了SpeedrunBench,这是一个旨在评估大型语言模型代理策略形成能力的新基准。该基准使用九款不同游戏的视频游戏速通,挑战代理改进其策略并超越之前的尝试。虽然目前的前沿代理在更简单的游戏中显示出潜力,但在实际计算预算内,它们在更复杂、持续时间更长的游戏中仍落后于人类表现。 AI

影响 该基准可以推动大型语言模型代理策略形成和长时推理能力的进步。

排序理由 该项目描述了一个用于评估大型语言模型代理的新基准,该项目发表在arXiv的学术论文中。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的SpeedrunBench用视频游戏速通挑战大型语言模型代理

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个用于评估大型语言模型代理的新基准,该项目发表在arXiv的学术论文中。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yoshinari Fujinuma, Keisuke Kamahori, Ryuto Koike, Abdelrahman Madkour, Varun Prashant Gangal, Monty Bichouna, Martyna Markiewicz, Shivani Jain, Duncan Curtis, Rebecca Qian, Anand Kannappan ·

    SpeedrunBench:用视频游戏速通挑战大型语言模型代理

    arXiv:2610.08076v1 Announce Type: new Abstract: Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved…