PulseAugur
EN
LIVE 23:11:07

New models enhance robot manipulation by integrating vision and state

Researchers have developed several new methods to improve robot manipulation capabilities by better integrating visual information with the robot's state and actions. GeoProp, for instance, is a lightweight adapter that aligns proprioception with vision by projecting robot state onto image planes and injecting spatial priors. RoboDojo offers a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks. DSWAM introduces a dual-system approach combining a World Action Model executor with a vision-language planner for fine-grained manipulation, while DynaWM uses a base-VLA-guided world foundation model specifically for dynamic object manipulation. AI

IMPACT These advancements aim to improve robot dexterity and adaptability in complex, real-world scenarios, potentially accelerating the deployment of more capable robotic systems.

RANK_REASON Multiple research papers introducing new models and benchmarks for robot manipulation.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 9 sources. How we write summaries →

New models enhance robot manipulation by integrating vision and state

COVERAGE [9]

  1. arXiv cs.AI TIER_1 English(EN) · Guoyang Zhao, Quanhao Qian, Gongjie Zhang, Wenhao Li, Jiuniu Wang, Xiaowei Lu, Deli Zhao, Ran Xu ·

    GeoProp: Grounding Robot State in Vision for Generalist Manipulation

    arXiv:2607.07101v1 Announce Type: cross Abstract: Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D fe…

  2. arXiv cs.AI TIER_1 English(EN) · Ran Xu ·

    GeoProp: Grounding Robot State in Vision for Generalist Manipulation

    Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to grou…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    GeoProp: Grounding Robot State in Vision for Generalist Manipulation

    Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to grou…

  4. arXiv cs.AI TIER_1 English(EN) · Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Weijie Wan, Baijun Chen, Haoran Lu, Haowen Yan, Honghao Su, Zhiyang Dou, Kaixuan Wang, Dandan Zhang, Yunze Liu, Yan Qin, Qiwei Liang, Qiwei Wu, Zijian Lin, Wenwei Lin, Yuran Wang, Minghua He, T… ·

    RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

    arXiv:2607.04434v1 Announce Type: cross Abstract: Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capabilit…

  5. arXiv cs.AI TIER_1 English(EN) · Jian Zhu, Jianjun Zhang, Taiyi Su, Tianbin Liu, Zhangyuan Wang, Kai Xie, Zitai Huang, Chong Ma, Youzhang He, Tianjian Wang, Hanyang Wang, Weihao Ding, Yi Xu ·

    DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

    arXiv:2607.04927v1 Announce Type: cross Abstract: World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded exe…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

    A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics policy learning.

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

    RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.

  8. arXiv cs.AI TIER_1 English(EN) · Yi Xu ·

    DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

    World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-l…

  9. arXiv cs.CV TIER_1 English(EN) · Chongkei Chang, Zhidong Deng ·

    DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation

    arXiv:2607.02604v1 Announce Type: new Abstract: Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-end forward or inverse dynamics models, i.e., world …