PulseAugur
实时 14:50:54
English(EN) Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

Action QFormer 通过塑造多模态表示来增强 VLA 模型

研究人员推出 Action QFormer,一种通过将动作监督视为表示塑造机制而非仅仅下游目标来增强视觉-语言-动作 (VLA) 模型的新方法。该方法利用指令条件查询来重组多模态信息,创建与动作兼容的表示,而不会破坏语言处理或对象基础。在零样本 sim-to-real 导航任务中,Action QFormer 将闭环任务成功率从 18.8% 提高到 56.3%,将动作生成正确率从 22.5% 提高到 75.5%,同时还减少了分布外生成。 AI

影响 这项研究可能为机器人和具身人工智能等任务带来更强大、更准确的 VLA 模型。

排序理由 该集群包含一篇详细介绍 AI 模型新方法的论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Action QFormer 通过塑造多模态表示来增强 VLA 模型

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Yufeng Ji, Wenhao Tang, Haoyi Niu, Koushil Sreenath, Yi Wu, Zhongyu Li ·

    Action QFormer:在视觉-语言-动作模型中通过动作监督进行结构化表示塑造

    arXiv:2607.14635v1 Announce Type: new Abstract: Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We …

  2. arXiv cs.CV TIER_1 English(EN) · Zhongyu Li ·

    Action QFormer:在视觉-语言-动作模型中通过动作监督进行结构化表示塑造

    Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is …