PulseAugur
EN
LIVE 13:18:28

New VLA models enhance robot manipulation with foresight and large-scale data

Researchers have developed new methods to improve the performance of Vision-Language-Action (VLA) models in robotics, particularly for complex, long-horizon tasks. One approach, Foresight Residual RL, enhances credit assignment by incorporating a foresight value that predicts future subtask success, leading to significantly higher overall task completion rates. Another development, Xiaomi-Robotics-1, scales VLA models using over 100,000 hours of real-world robot trajectories and an auto-labeling pipeline, demonstrating strong performance and efficient fine-tuning capabilities on various benchmarks. Additionally, a technique using robot-centric pointmaps addresses frame mismatches in VLA models, improving generalization across different camera viewpoints and robot embodiments. AI

IMPACT These advancements in VLA models and training methodologies are expected to accelerate the development of more capable and generalizable robots for complex manipulation tasks.

RANK_REASON Multiple research papers detailing novel methods and models for Vision-Language-Action (VLA) in robotics.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 7 sources. How we write summaries →

New VLA models enhance robot manipulation with foresight and large-scale data

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple research papers detailing novel methods and models for Vision-Language-Action (VLA) in robotics.
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
71 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [7]

  1. arXiv cs.AI TIER_1 English(EN) · Jin Wang, Kim Tien Ly, Jacques Cloete, Jin Jin, Nikos Tsagarakis, Ioannis Havoutis ·

    Vision-Language-Policy Model for Dynamic Robot Task Planning

    arXiv:2512.19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene throu…

  2. arXiv cs.AI TIER_1 English(EN) · Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov ·

    Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

    arXiv:2607.19399v1 Announce Type: cross Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-di…

  3. arXiv cs.AI TIER_1 English(EN) · Tingzheng Jia, Kan Guo, Lanping Qian, Yongli Hu, Daxin Tian, Guixian Qu, Chunmian Lin, Baocai Yin, Jiapu Wang ·

    AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models

    arXiv:2604.17787v2 Announce Type: replace-cross Abstract: Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic fo…

  4. arXiv cs.LG TIER_1 English(EN) · Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias ·

    Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models

    arXiv:2607.16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometric…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream task…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  7. arXiv cs.CV TIER_1 English(EN) · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, … ·

    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    arXiv:2607.15330v1 Announce Type: cross Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and…