English(EN)ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
新的世界-动作模型增强机器人操作和泛化能力
作者PulseAugur 编辑部·[8 个来源]·
研究人员开发了几种新的机器人操作世界-动作模型(WAMs),旨在提高效率和鲁棒性。LiLa-WAM专注于轻量级潜在推理空间,可在单个GPU上进行端到端训练,在基准任务上取得了很高的成功率。ST-WAM通过使用具有DINOv3特征和历史检索的语义-时间建模,增强了在视觉分布变化下的鲁棒性,在真实世界场景中表现明显优于先前模型。LAWM-3D从人类视频中学习3D感知潜在动作,通过多视图不变性和几何对齐来提高世界模型性能。XEWorld引入了一个测试平台,用于评估对未见过机器人实体的泛化能力,揭示了当前模型主要充当2D视觉模式匹配器。最后,MobileWAM通过融合视频扩散Transformer与动作专家并采用“先见之明链”(Chain-of-Foresight)方法进行运动和操作,将WAMs应用于移动操作。
AI
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recent…
arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial comp…
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limitin…
arXiv:2608.05706v1 Announce Type: new Abstract: World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterog…
arXiv:2608.05799v1 Announce Type: cross Abstract: Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual pattern…
arXiv:2608.04657v1 Announce Type: new Abstract: World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation ami…
arXiv cs.CV
TIER_1English(EN)·Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li·
arXiv:2607.28993v1 Announce Type: cross Abstract: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transi…
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet e…