English(EN)Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
新的VLA模型通过远见和大规模数据增强机器人操作能力
作者PulseAugur 编辑部·[7 个来源]·
研究人员开发了新的方法来提高机器人领域中视觉-语言-动作(VLA)模型的性能,特别是在复杂、长周期的任务方面。一种方法是“远见残差强化学习”(Foresight Residual RL),它通过引入预测未来子任务成功率的远见值来增强信用分配,从而显著提高整体任务完成率。另一项开发是Xiaomi-Robotics-1,它使用超过10万小时的真实世界机器人轨迹和一个自动标注流程来扩展VLA模型,在各种基准测试中展示了强大的性能和高效的微调能力。此外,一种使用以机器人为中心的点图(pointmaps)的技术解决了VLA模型中的帧不匹配问题,提高了在不同摄像头视角和机器人实体上的泛化能力。
AI
arXiv:2512.19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene throu…
arXiv cs.AI
TIER_1English(EN)·Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov·
arXiv:2607.19399v1 Announce Type: cross Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-di…
arXiv:2604.17787v2 Announce Type: replace-cross Abstract: Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic fo…
arXiv:2607.16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometric…
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream task…
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…
arXiv cs.CV
TIER_1English(EN)·Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, …·
arXiv:2607.15330v1 Announce Type: cross Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and…