English(EN)World Action Learning via Interaction-Centric Spectral Latent Guidance
新方法通过视频数据推进机器人策略学习 · 跟踪4个来源
作者PulseAugur 编辑部·[6 个来源]·
研究人员开发了训练世界动作模型的新方法,该模型可以预测未来的视觉动态和机器人动作。一种方法 RLA World Model (RLA-WM) 利用源自 DINO 残差的新型残差潜在动作 (RLA) 表示,在效率和预测准确性方面优于现有的基于特征和视频扩散模型。另一种方法 NAVA-WAM 通过直接从仅观察视频预训练动作策略来引入原生动作先验学习,在各种设置中表现出强大的性能和泛化能力。第三个框架 WING 专注于通过分离观察者引起的运动与手部物体交互,并使用谱引导来对齐人类和机器人行为,从而将交互知识从以自我为中心的视频转移到机器人策略。
AI
arXiv:2605.07079v2 Announce Type: replace-cross Abstract: World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features inste…
World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typicall…
arXiv:2610.09427v1 Announce Type: new Abstract: World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, w…
arXiv:2610.09514v1 Announce Type: new Abstract: Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We const…
arXiv cs.CV
TIER_1English(EN)·Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost, Zijian Zhou, Yikai Wang, Xudong Wang, Aditya Patel, Belinda Zeng, Tao Xiang, Serge Belongie, Amir Bar, Sen He·
arXiv:2610.03391v1 Announce Type: new Abstract: World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about intera…