English(EN)Geometric Action Model for Robot Policy Learning
新的机器人策略模型增强了动作生成和效率
作者PulseAugur 编辑部·[7 个来源]·
研究人员开发了新的机器人策略学习方法,提高了动作生成效率和准确性。LeaP(一种可学习的源先验)通过对本体感觉进行条件化来优化动作生成的起点,从而在操作任务上取得了显著的性能提升。LaWAM引入了潜在世界动作模型,该模型预测紧凑的潜在视觉子目标而非完整的视频帧,从而在保持高成功率的同时降低了计算延迟。几何动作模型(GAM)将几何基础模型重新用于语言条件操作,直接整合3D几何以实现更鲁棒、更高效的控制。
AI
arXiv:2606.18594v1 Announce Type: cross Abstract: In real-world reinforcement learning (RL), the choice of action space can play a key role in shaping motion smoothness, safety, and overall task performance. In this study, we evaluate pose increment, pose velocity, joint position…
arXiv:2606.17408v1 Announce Type: cross Abstract: Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored. This work asks a simple question: where should ac…
arXiv:2606.15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this l…
arXiv cs.LG
TIER_1English(EN)·Jisang Han, Seonghu Jeon, Jaewoo Jung, Ren\'e Zurbr\"ugg, Honggyu An, Tifanny Portela, Marco Hutter, Marc Pollefeys, Seungryong Kim, Sunghwan Hong·
arXiv:2606.17046v1 Announce Type: cross Abstract: Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action models (WA…
A geometric action model leverages pretrained geometric foundation models to enable language-conditioned manipulation policies with improved accuracy, robustness, and efficiency in 3D physical environments.
LaWAM enables efficient robot control by predicting compact latent visual subgoals instead of expensive video generation, achieving high performance with reduced computational latency.
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action models (WAMs) inherit strong semantic or temporal priors fro…