PulseAugur
实时 21:56:16
English(EN) Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

新的VLA模型通过远见和大规模数据增强机器人操作能力

研究人员开发了新的方法来提高机器人领域中视觉-语言-动作(VLA)模型的性能,特别是在复杂、长周期的任务方面。一种方法是“远见残差强化学习”(Foresight Residual RL),它通过引入预测未来子任务成功率的远见值来增强信用分配,从而显著提高整体任务完成率。另一项开发是Xiaomi-Robotics-1,它使用超过10万小时的真实世界机器人轨迹和一个自动标注流程来扩展VLA模型,在各种基准测试中展示了强大的性能和高效的微调能力。此外,一种使用以机器人为中心的点图(pointmaps)的技术解决了VLA模型中的帧不匹配问题,提高了在不同摄像头视角和机器人实体上的泛化能力。 AI

影响 这些VLA模型和训练方法学的进步预计将加速开发更强大、更通用的机器人,以应对复杂的操纵任务。

排序理由 多篇研究论文详细介绍了机器人领域中视觉-语言-动作(VLA)的新方法和模型。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新的VLA模型通过远见和大规模数据增强机器人操作能力

报道来源 [4]

  1. arXiv cs.LG TIER_1 English(EN) · Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias ·

    面向长时域机器人操作的视觉-语言-动作模型的前瞻残差强化学习

    arXiv:2607.16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometric…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Xiaomi-Robotics-1:使用超过10万小时的真实世界轨迹扩展视觉-语言-动作模型

    We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream task…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  4. arXiv cs.CV TIER_1 English(EN) · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, … ·

    Xiaomi-Robotics-1:使用超过10万小时的真实世界轨迹扩展视觉-语言-动作模型

    arXiv:2607.15330v1 Announce Type: cross Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and…