PulseAugur
EN
LIVE 20:31:05

New VLA models enhance robot manipulation with foresight and large-scale data

Researchers have developed new methods to improve the performance of Vision-Language-Action (VLA) models in robotics, particularly for complex, long-horizon tasks. One approach, Foresight Residual RL, enhances credit assignment by incorporating a foresight value that predicts future subtask success, leading to significantly higher overall task completion rates. Another development, Xiaomi-Robotics-1, scales VLA models using over 100,000 hours of real-world robot trajectories and an auto-labeling pipeline, demonstrating strong performance and efficient fine-tuning capabilities on various benchmarks. Additionally, a technique using robot-centric pointmaps addresses frame mismatches in VLA models, improving generalization across different camera viewpoints and robot embodiments. AI

IMPACT These advancements in VLA models and training methodologies are expected to accelerate the development of more capable and generalizable robots for complex manipulation tasks.

RANK_REASON Multiple research papers detailing novel methods and models for Vision-Language-Action (VLA) in robotics.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New VLA models enhance robot manipulation with foresight and large-scale data

COVERAGE [4]

  1. arXiv cs.LG TIER_1 English(EN) · Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias ·

    Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models

    arXiv:2607.16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometric…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream task…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  4. arXiv cs.CV TIER_1 English(EN) · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, … ·

    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    arXiv:2607.15330v1 Announce Type: cross Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and…