PulseAugur
实时 12:04:13
English(EN) Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

新的VLA模型通过远见和大规模数据增强机器人操作能力

研究人员开发了新的方法来提高机器人领域中视觉-语言-动作(VLA)模型的性能,特别是在复杂、长周期的任务方面。一种方法是“远见残差强化学习”(Foresight Residual RL),它通过引入预测未来子任务成功率的远见值来增强信用分配,从而显著提高整体任务完成率。另一项开发是Xiaomi-Robotics-1,它使用超过10万小时的真实世界机器人轨迹和一个自动标注流程来扩展VLA模型,在各种基准测试中展示了强大的性能和高效的微调能力。此外,一种使用以机器人为中心的点图(pointmaps)的技术解决了VLA模型中的帧不匹配问题,提高了在不同摄像头视角和机器人实体上的泛化能力。 AI

影响 这些VLA模型和训练方法学的进步预计将加速开发更强大、更通用的机器人,以应对复杂的操纵任务。

排序理由 多篇研究论文详细介绍了机器人领域中视觉-语言-动作(VLA)的新方法和模型。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

新的VLA模型通过远见和大规模数据增强机器人操作能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文详细介绍了机器人领域中视觉-语言-动作(VLA)的新方法和模型。
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
71 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [7]

  1. arXiv cs.AI TIER_1 English(EN) · Jin Wang, Kim Tien Ly, Jacques Cloete, Jin Jin, Nikos Tsagarakis, Ioannis Havoutis ·

    用于动态机器人任务规划的视觉-语言-策略模型

    arXiv:2512.19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene throu…

  2. arXiv cs.AI TIER_1 English(EN) · Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov ·

    利用离线监督实现大规模视觉-语言-动作模型的高效和可泛化强化学习

    arXiv:2607.19399v1 Announce Type: cross Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-di…

  3. arXiv cs.AI TIER_1 English(EN) · Tingzheng Jia, Kan Guo, Lanping Qian, Yongli Hu, Daxin Tian, Guixian Qu, Chunmian Lin, Baocai Yin, Jiapu Wang ·

    AnchorRefine:基于轨迹锚点和残差精炼的视觉-语言-动作模型协同操控

    arXiv:2604.17787v2 Announce Type: replace-cross Abstract: Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic fo…

  4. arXiv cs.LG TIER_1 English(EN) · Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias ·

    面向长时域机器人操作的视觉-语言-动作模型的前瞻残差强化学习

    arXiv:2607.16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometric…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    Xiaomi-Robotics-1:使用超过10万小时的真实世界轨迹扩展视觉-语言-动作模型

    We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream task…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    像机器人一样看:面向视觉-语言-动作模型的机器人中心点图

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  7. arXiv cs.CV TIER_1 English(EN) · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, … ·

    Xiaomi-Robotics-1:使用超过10万小时的真实世界轨迹扩展视觉-语言-动作模型

    arXiv:2607.15330v1 Announce Type: cross Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and…