PulseAugur
实时 01:07:17
English(EN) DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

新模型通过整合视觉和状态来增强机器人操作能力

研究人员开发了几种新方法,通过更好地整合视觉信息与机器人的状态和动作来提高机器人操作能力。例如,GeoProp 是一种轻量级适配器,通过将机器人状态投影到图像平面并注入空间先验来对齐本体感觉和视觉。RoboDojo 提供了一个统一的模拟和真实基准,用于评估通用机器人操作策略在各种任务中的表现。DSWAM 引入了一种双系统方法,将世界动作模型执行器与视觉语言规划器相结合,以实现细粒度操作,而 DynaWM 使用专门针对动态物体操作的基于 VLA 的世界基础模型。 AI

影响 这些进步旨在提高机器人在复杂现实场景中的灵活性和适应性,有可能加速部署更强大的机器人系统。

排序理由 多篇研究论文介绍了用于机器人操作的新模型和基准。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

新模型通过整合视觉和状态来增强机器人操作能力

报道来源 [9]

  1. arXiv cs.AI TIER_1 English(EN) · Guoyang Zhao, Quanhao Qian, Gongjie Zhang, Wenhao Li, Jiuniu Wang, Xiaowei Lu, Deli Zhao, Ran Xu ·

    GeoProp:在视觉中为通用机器人操纵奠定机器人状态基础

    arXiv:2607.07101v1 Announce Type: cross Abstract: Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D fe…

  2. arXiv cs.AI TIER_1 English(EN) · Ran Xu ·

    GeoProp: 在视觉中为通用机器人操控打下机器人状态基础

    Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to grou…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    GeoProp:在视觉中为通用机器人操纵奠定机器人状态基础

    Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to grou…

  4. arXiv cs.AI TIER_1 English(EN) · Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Weijie Wan, Baijun Chen, Haoran Lu, Haowen Yan, Honghao Su, Zhiyang Dou, Kaixuan Wang, Dandan Zhang, Yunze Liu, Yan Qin, Qiwei Liang, Qiwei Wu, Zijian Lin, Wenwei Lin, Yuran Wang, Minghua He, T… ·

    RoboDojo:一个统一的仿真与真实基准,用于全面评估通用机器人操作策略

    arXiv:2607.04434v1 Announce Type: cross Abstract: Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capabilit…

  5. arXiv cs.AI TIER_1 English(EN) · Jian Zhu, Jianjun Zhang, Taiyi Su, Tianbin Liu, Zhangyuan Wang, Kai Xie, Zitai Huang, Chong Ma, Youzhang He, Tianjian Wang, Hanyang Wang, Weihao Ding, Yi Xu ·

    DSWAM:用于细粒度机器人操作的双系统世界动作基础模型

    arXiv:2607.04927v1 Announce Type: cross Abstract: World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded exe…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    RynnWorld-4D: 用于机器人操作的4D具身世界模型

    A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics policy learning.

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboDojo:一个统一的仿真与真实基准,用于全面评估通用机器人操控策略

    RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.

  8. arXiv cs.AI TIER_1 English(EN) · Yi Xu ·

    DSWAM:用于细粒度机器人操作的双系统世界动作基础模型

    World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-l…

  9. arXiv cs.CV TIER_1 English(EN) · Chongkei Chang, Zhidong Deng ·

    DynaWM:一个基于VLA的动态物体操控世界基础模型

    arXiv:2607.02604v1 Announce Type: new Abstract: Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-end forward or inverse dynamics models, i.e., world …