PulseAugur
中
实时 13:28:20
English(EN) ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

新的世界-动作模型增强机器人操作和泛化能力

研究人员开发了几种新的机器人操作世界-动作模型(WAMs),旨在提高效率和鲁棒性。LiLa-WAM专注于轻量级潜在推理空间,可在单个GPU上进行端到端训练,在基准任务上取得了很高的成功率。ST-WAM通过使用具有DINOv3特征和历史检索的语义-时间建模,增强了在视觉分布变化下的鲁棒性,在真实世界场景中表现明显优于先前模型。LAWM-3D从人类视频中学习3D感知潜在动作,通过多视图不变性和几何对齐来提高世界模型性能。XEWorld引入了一个测试平台,用于评估对未见过机器人实体的泛化能力,揭示了当前模型主要充当2D视觉模式匹配器。最后,MobileWAM通过融合视频扩散Transformer与动作专家并采用“先见之明链”(Chain-of-Foresight)方法进行运动和操作,将WAMs应用于移动操作。 AI

影响 这些世界-动作模型的进步可能导致机器人在复杂、真实世界的环境中更具能力和适应性。

排序理由 多篇研究论文介绍了用于机器人操作的新模型和框架。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新的世界-动作模型增强机器人操作和泛化能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了用于机器人操作的新模型和框架。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
69 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [8]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    LAWM-3D:从人类视频中学习 3D 感知潜在动作,用于可泛化的机器人世界模型

    World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recent…

  2. arXiv cs.AI TIER_1 English(EN) · Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You, Fugui Fan, Yuting Wu, Minghui Wu, Chenxu Zhao, JiaHong Ning, Peiguang Jing ·

    LiLa-WAM:轻量级潜在推理世界-动作模型用于机器人操作

    arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial comp…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    ST-WAM:视觉分布变化下的鲁棒操纵语义-时序世界动作模型

    World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limitin…

  4. arXiv cs.CV TIER_1 English(EN) · Jiarui Yang, Jiale Zhange, Jiawei Li, Hang Guo, Wen Huang, Jinpeng Wang, Peidong Liu, Shu-Tao Xia ·

    LAWM-3D:从人类视频中学习 3D 感知潜在动作,用于可泛化的机器人世界模型

    arXiv:2608.05706v1 Announce Type: new Abstract: World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterog…

  5. arXiv cs.CV TIER_1 English(EN) · Yixiang Chen, Jiabing Yang, Yuan Xu, Qisen Ma, Keji He, Peiyan Li, Kai Wang, Ziheng He, Xiangnan Wu, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang ·

    XEWorld:动作条件世界模型能否泛化到未见的机器人实体?

    arXiv:2608.05799v1 Announce Type: cross Abstract: Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual pattern…

  6. arXiv cs.CV TIER_1 English(EN) · Zehua Fan, Junjie He, Wenxuan Song, Xi Wang, Wenqi Lyu, Linge Zhao, Fuhao Li, Zihan You, Yifei Yang, Kaiming Xu, Qi Jiang, Yue Jiang, Haoang Li, Cheng Chi, Bailin Li, Yan Wang ·

    MobileWAM:将世界动作模型桥接到具有远见链的移动操作

    arXiv:2608.04657v1 Announce Type: new Abstract: World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation ami…

  7. arXiv cs.CV TIER_1 English(EN) · Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li ·

    ST-WAM:视觉分布变化下的鲁棒操纵语义-时序世界动作模型

    arXiv:2607.28993v1 Announce Type: cross Abstract: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transi…

  8. Mastodon — mastodon.social TIER_1 English(EN) · iankhanfuturist ·

    $ω$-0:一个潜在的预测性世界动作模型,用于并发人形移动操作

    $ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet e…