PulseAugur
中
实时 10:30:24
English(EN) World2Motion: Turning Video World Models into 3D Human Motion Generators

新研究增强了机器人和AI的世界动作模型

近期研究探索了用于机器人和AI的世界动作模型(WAMs)的进展,重点是提高预测精度、动作生成和推理效率。几篇论文介绍了新的方法,如完成感知引导(CAG)以确保任务完成,回顾性世界建模以实现向后推理,以及动作经验字典(AED)以实现技能重用。其他工作通过诸如动作引导稀疏想象(Sparse-WAM)和带有阶梯策略的流式推理等技术来解决计算成本问题。这些创新旨在提高机器人控制、泛化能力和在复杂环境中的鲁棒性。 AI

影响 这些进展旨在通过增强世界建模能力来改进机器人控制、规划和泛化。

排序理由 多篇arXiv论文发表,详细介绍了世界动作模型的新方法和评估。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 29 个来源。 我们如何撰写摘要 →

新研究增强了机器人和AI的世界动作模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇arXiv论文发表,详细介绍了世界动作模型的新方法和评估。
Source corroboration
29 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
10 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [29]

  1. arXiv cs.AI TIER_1 English(EN) · Seungyeon Kim, Junhoo Lee, Baekseung Kim, Minkyu Kim, Nojun Kwak ·

    Completion Aware Guidance for World Action Models

    arXiv:2610.01559v1 Announce Type: cross Abstract: World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In thi…

  2. arXiv cs.AI TIER_1 English(EN) · Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar ·

    用于世界建模的潜在视频预测:一项揭示引人入胜的有利证据的评估

    arXiv:2605.15618v2 Announce Type: replace-cross Abstract: Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reported as a final task score, obscuring how their representations behave under the deg…

  3. arXiv cs.AI TIER_1 English(EN) · Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan ·

    从历史动作轨迹中学习技能:世界动作模型的动作经验词典

    arXiv:2609.40219v1 Announce Type: cross Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture under…

  4. arXiv cs.AI TIER_1 English(EN) · Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo ·

    超越预测:通过回顾性世界模型引导VLM智能体

    arXiv:2609.39101v1 Announce Type: new Abstract: Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods m…

  5. arXiv cs.AI TIER_1 English(EN) · Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo ·

    Sparse-WAM:通过动作引导的稀疏想象加速世界动作模型

    arXiv:2609.38984v1 Announce Type: cross Abstract: World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    World Observer:用于持久化世界建模的联合 Actor-Observer 生成

    How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often …

  7. arXiv cs.AI TIER_1 English(EN) · Guoheng Sun, Chen Chen, Jin Wang, Ang Li, Teresa Lv ·

    Staircase Policy: 具有大动作块的世界-动作模型的流式推理

    arXiv:2609.36471v1 Announce Type: cross Abstract: World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Ac…

  8. arXiv cs.AI TIER_1 English(EN) · Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li ·

    V-JEPA策略:在预测性视觉潜在空间上构建有效的世界-动作模型

    arXiv:2609.37250v1 Announce Type: cross Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge …

  9. arXiv cs.AI TIER_1 English(EN) · Cunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou, Yue Yao, Chi Harold Liu ·

    AquaWAM:面向水下具身智能体的动态感知世界动作模型

    arXiv:2609.33299v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environmen…

  10. arXiv cs.LG TIER_1 English(EN) · Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu ·

    FlexiWorld:跨越多个时间尺度的灵活动作块学习与规划

    arXiv:2609.35138v2 Announce Type: replace Abstract: Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or…

  11. arXiv cs.AI TIER_1 English(EN) · Xiangcheng Zhan, Zirui Chen, Yicheng Zhao, Ziteng Gao, Shuo Yang ·

    直接体验世界模型优化:超越动作模仿学习世界

    arXiv:2609.37398v1 Announce Type: new Abstract: World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Esp…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    Memorizon:训练超出其上下文窗口的世界模型

    Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervi…

  13. arXiv cs.AI TIER_1 English(EN) · Ke He, Yichen Ding, Bin Yang ·

    OneWorld:在世界模型中学习跨动作的一致物理学

    arXiv:2609.30946v1 Announce Type: cross Abstract: Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generat…

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    EVO-WAM:通过视频-动作验证实现世界动作模型的演进

    Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for …

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    WorldPlay2:在控制与视野中扩展实时交互式世界模型

    Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency a…

  16. Hugging Face Daily Papers TIER_1 English(EN) ·

    AnyStep-WAM:面向世界动作模型的预算对齐蒸馏与自适应推理

    World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require p…

  17. Hugging Face Daily Papers TIER_1 English(EN) ·

    InternW0-Δ:一个连接预测动力学与动作的世界动作模型,拥有超过20K小时的开放数据

    World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified fram…

  18. arXiv cs.CV TIER_1 English(EN) · Xindi Yang, Baolu Li, Liam Lee, Zhenfei Yin, Songxin Zhang, Zhuoyang Song, Xu Jia, Jianfei Cai, Tien-Tsin Wong, Bingyi Jing, Mengyue Yang ·

    Oneira:从开放式生成到视频世界模型中的开放世界交互

    arXiv:2610.01614v1 Announce Type: new Abstract: Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly creat…

  19. arXiv cs.CV TIER_1 English(EN) · Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang ·

    FutureWorlds:从替代未来学习机器人世界模型

    arXiv:2610.01019v1 Announce Type: new Abstract: Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates l…

  20. arXiv cs.CV TIER_1 English(EN) · Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen ·

    CtrlWAM:具有对齐意图和远见的控制世界动作模型

    arXiv:2610.00859v1 Announce Type: new Abstract: World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed action…

  21. arXiv cs.CV TIER_1 English(EN) · Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin, Jinhyeok Choi, Junyoung Seo, Seungryong Kim ·

    World Observer:用于持久化世界建模的联合行动者-观察者生成

    arXiv:2610.02162v1 Announce Type: new Abstract: How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, th…

  22. arXiv cs.CV TIER_1 English(EN) · Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu ·

    Memorizon:训练超越其上下文窗口的世界模型

    arXiv:2610.00544v1 Announce Type: new Abstract: Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span …

  23. arXiv cs.CV TIER_1 English(EN) · Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang ·

    CST-WM:一种用于具身视觉跟踪的因果结构化世界模型

    arXiv:2609.06302v2 Announce Type: replace Abstract: Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning…

  24. arXiv cs.CV TIER_1 English(EN) · Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of… ·

    事件边界的远见:评估视频世界模型中的物理预测

    arXiv:2609.36531v1 Announce Type: new Abstract: Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physica…

  25. arXiv cs.CV TIER_1 English(EN) · Tu Fangyuan, Xiangyue Zhang, Yiyi Cai, Yichen Peng, Kunhang Li, Bo Zheng, Zhixiang Wang, Kaipeng Zhang, Erwin Wu, Haoran Xie, Haiyang Liu ·

    World2Motion:将视频世界模型转化为3D人体运动生成器

    arXiv:2609.37004v1 Announce Type: new Abstract: We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is c…

  26. arXiv cs.CV TIER_1 English(EN) · Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian ·

    EVO-WAM:通过视频-动作验证实现世界动作模型的演进

    arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions,…

  27. arXiv cs.CV TIER_1 English(EN) · Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang ·

    重新思考世界-动作建模的表征

    arXiv:2609.38163v1 Announce Type: new Abstract: World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding…

  28. arXiv cs.CV TIER_1 English(EN) · Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang, Taojun Ding, Tingyang Xiao, Zhizhong Su, Jie Wang, Sanping Zhou ·

    CogWAM:通过事件驱动接口将语义认知与世界行动建模对齐

    arXiv:2609.37721v1 Announce Type: cross Abstract: Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM,…

  29. arXiv cs.CV TIER_1 English(EN) · Xueji Fang, Boqiang Duan, Hua Wu, Jingdong Wang, Guo-Jun Qi ·

    潜在演化世界行动模型

    arXiv:2609.27455v2 Announce Type: replace Abstract: World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting …