PulseAugur
中
实时 08:55:30
English(EN) Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA

在模拟MOBA世界中训练的AI代理在真实游戏中获胜率达到70%

研究人员开发了一种新颖的方法,通过利用连续的Dyna循环来训练AI代理以应对复杂游戏。该方法仅在十人英雄多人在线战术竞技游戏(MOBA)的已学习世界模型中训练策略,真实游戏仅用于为世界模型更新和策略评估提供数据。训练好的策略在实际游戏中获胜率达到70.2%,与仅在想象中训练的零胜率相比有了显著提高。关键发现表明,模型利用难以在模拟环境中进行评估,并且在策略自身游戏场景下,在训练数据上准确的世界模型可能不准确,因此需要Dyna循环进行修复。 AI

影响 展示了一种在复杂环境中训练AI代理的新颖方法,有可能加速AI在策略游戏及其他领域的开发。

排序理由 详细介绍新AI训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

在模拟MOBA世界中训练的AI代理在真实游戏中获胜率达到70%

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍新AI训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jordy Kieto ·

    梦中学习,现实制胜:十人MOBA的连续动力循环

    arXiv:2610.08033v1 Announce Type: new Abstract: World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model …