PulseAugur
实时 14:03:56
中文(ZH) 把 Gemini 和 GPT 放到《我的世界》当教练,谁能看懂机器人的微操?

清华大学与英伟达的具身AI评估:Gemini领先,GPT落后

来自国立清华大学和英伟达的研究人员开发了一个名为VLM-AR3L的新框架,用于评估视觉语言模型(VLMs)在具身智能任务中的性能。在他们的研究中,Gemini 2.0在各种任务中表现强劲,而GPT-4.1则显示出局限性。该框架通过使用VLMs进行离线标注,并训练更小、更轻量级的网络以实现实时指导,从而解决了高API成本和延迟的挑战。VLM-AR3L结合了绝对奖励和相对奖励,为强化学习代理提供全局方向和细粒度进度评估。 AI

影响 这项研究为评估和训练具身AI代理提供了一个新框架,有望加速机器人技术和交互式AI系统的进展。

排序理由 研究论文,详细介绍了新框架以及对现有VLMs在具身任务上的评估。[lever_c_demoted from research: ic=1 ai=1.0]

在 雷峰网 (Leiphone) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

清华大学与英伟达的具身AI评估:Gemini领先,GPT落后

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
研究论文,详细介绍了新框架以及对现有VLMs在具身任务上的评估。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    将Gemini和GPT引入Minecraft担任教练,谁能理解机器人的微操?

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><img class="rich_pages wxw-img" src="https://static.leiphone.com/uploads/new/images/20260903/6a9946fa4e5e2.jpg?imageMogr2/quality/90" style="width: 100%; display: inline-block; text-align:…