PulseAugur
中
实时 16:08:37
中文(ZH) 具身智能“高考”难疯了!人类100分,最强模型12.8

RoboDojo基准揭示AI机器人与人类表现的巨大差距

一项名为RoboDojo的新基准已发布,用于评估具身智能,包含42个模拟任务和18个真实世界机器人任务。该基准突显了当前AI模型与人类表现之间存在的巨大差距,在模拟环境中,最佳模型成功率仅为8.80%,在真实机器人上为12.8%,而人类的成功率分别为76.03%和100%。RoboDojo旨在为具身智能提供标准化和全面的评估,涵盖泛化能力、记忆、精度、长时序执行和开放式语义理解。 AI

影响 凸显了当前AI在现实世界机器人领域能力的显著不足,表明需要更强大的模型。

排序理由 该集群描述了一个新的具身智能基准和评估框架,包括模型性能的研究发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 量子位 (QbitAI) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

RoboDojo基准揭示AI机器人与人类表现的巨大差距

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个新的具身智能基准和评估框架,包括模型性能的研究发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
91 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    具身智能“高考”引爆热潮!人类考100分,最强模型仅得12.8分

    具身测评界的珠峰来了:RoboDojo