PulseAugur
实时 10:25:36
English(EN) SceneActBench: Can Agents Act on the 3D Scenes They See?

新的SceneActBench基准评估VLM智能体在3D交互方面的能力

研究人员推出了SceneActBench,一个旨在评估视觉语言模型(VLM)智能体在与3D环境交互和操作方面能力的新基准。与之前侧重于文本描述或单对象操作的基准不同,SceneActBench在一个统一的智能体-环境循环中,评估了智能体在五个不同3D任务中的表现。该基准使用PNG图像或视频帧,以及3D资产,来测试智能体在复杂、多对象3D场景中执行操作的能力。对十一种不同VLM配置的初步评估显示,性能范围在38.6-50.2之间,表明当前模型在所有任务上都难以持续表现出色。 AI

影响 该基准旨在提高VLM智能体在3D环境中交互和行动的能力,有望为涉及空间推理和操作的任务带来更强大的AI助手。

排序理由 该集群描述了一个用于评估AI模型的新基准,属于研究范畴。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的SceneActBench基准评估VLM智能体在3D交互方面的能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一个用于评估AI模型的新基准,属于研究范畴。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
54 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    SceneActBench:智能体能否对它们看到的3D场景进行操作?

    Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBe…

  2. arXiv cs.CV TIER_1 English(EN) · Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu ·

    SceneActBench:智能体能否对它们看到的3D场景采取行动?

    arXiv:2607.22393v1 Announce Type: cross Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-objec…