PulseAugur
EN
LIVE 19:50:58

New research advances world action models for robotic planning and control · 10 sources tracked

Recent research explores advancements in world action models (WAMs) for robotic control and planning. Several papers introduce new architectures and training methodologies to improve prediction accuracy, generalization, and efficiency. Key areas of focus include handling asynchronous execution, enhancing spatial understanding, and developing better methods for training and evaluating these models, particularly in complex, long-horizon tasks. AI

IMPACT These advancements could lead to more capable and generalizable robots, improving performance in complex manipulation and navigation tasks.

RANK_REASON Multiple arXiv papers published on related topics in world action models for robotics.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 22 sources. How we write summaries →

New research advances world action models for robotic planning and control · 10 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple arXiv papers published on related topics in world action models for robotics.
Source corroboration
22 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
6 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [22]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Sheng-Jun Huang ·

    Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

    World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the act…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

    World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the act…

  3. arXiv cs.AI TIER_1 English(EN) · Dhananjay Ashok, Shantanu Agarwal, Vivek Datla, Jonathan May, Alfy Samuel ·

    How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling

    arXiv:2610.02542v1 Announce Type: new Abstract: World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments, fine-tuning a Language Model (LM) to serve as a WM has emerged as a dom…

  4. arXiv cs.AI TIER_1 English(EN) · Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar ·

    World Action Modeling with Progressive Visual Planning

    arXiv:2610.02508v1 Announce Type: new Abstract: World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-hori…

  5. arXiv cs.AI TIER_1 English(EN) · Samuel Barbeau, Simon Roy, Giovanni Beltrame, Christian Desrosiers, Nicolas Thome ·

    Latent Goal Prediction from Language for Model-Based Planning

    arXiv:2606.20627v2 Announce Type: replace Abstract: Joint-Embedding Predictive Architectures (JEPAs) enable agents to plan in latent space by imagining the outcomes of candidate actions, yet task specification remains a bottleneck. Visual targets provide precise local gradients b…

  6. arXiv cs.AI TIER_1 English(EN) · Xiangcheng Zhang, Runhan Huang, Yilun Du ·

    World Action Planner: Generalizable Robot Decision-Making with Action-Conditioned World Models

    arXiv:2607.27599v2 Announce Type: replace Abstract: Building generalizable robot agents for diverse applications remains a fundamental challenge. While imitation learning-based policies can perform well in familiar training environments, they often struggle to generalize to novel…

  7. arXiv cs.LG TIER_1 English(EN) · Arjun Subramanian ·

    Counterfactual Action Evaluation, Observation Bottlenecks, and Representation Geometry in Joint-Embedding Predictive World Models

    arXiv:2610.02860v1 Announce Type: new Abstract: Low latent prediction error does not establish that a world model distinguishes the consequences of its actions. We introduce an evaluation protocol that traces the same intervention through simulator state, raster observations, tar…

  8. arXiv cs.LG TIER_1 English(EN) · Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li ·

    XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

    arXiv:2610.03516v1 Announce Type: cross Abstract: World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understandin…

  9. arXiv cs.LG TIER_1 English(EN) · Rohun Agrawal, Nimit Kalra, Arjun Parthasarathy, Yann LeCun, Oumayma Bounou, Pavel Izmailov, Micah Goldblum ·

    Closing the Train-Test Gap in World Models for Gradient-Based Planning

    arXiv:2512.09929v2 Announce Type: replace Abstract: World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning tasks at inference time. Compared to traditional…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    RealtimeWAM: One-Step Asynchronous World Action Models

    World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-…

  11. arXiv cs.LG TIER_1 English(EN) · Takumi Hara, Kanata Suzuki ·

    Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning

    arXiv:2610.01224v1 Announce Type: new Abstract: Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models w…

  12. arXiv cs.LG TIER_1 English(EN) · Zheyuan Zhang, Suyu Ye, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Daniel Khashabi, Tianmin Shu, Vaishnav Tadiparthi ·

    JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts

    arXiv:2610.00722v1 Announce Type: new Abstract: World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the lat…

  13. arXiv cs.LG TIER_1 English(EN) · Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke ·

    Learning Commute-Time-Preserving World Models for Planning

    arXiv:2610.01373v1 Announce Type: new Abstract: World models allow agents to plan in latent space by choosing a sequence of actions that most reduces the distance to a given goal state. Thus, planning can benefit from latent representations whose distances mirror commute-times in…

  14. arXiv cs.AI TIER_1 English(EN) · Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li, Letian Wang, Alexandre Alahi, Harold Soh ·

    Grounded World Model: Latent Planning with Language Goals

    arXiv:2604.11751v2 Announce Type: replace-cross Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning…

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    World Action Modeling with Progressive Visual Planning

    World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollou…

  16. Hugging Face Daily Papers TIER_1 English(EN) ·

    Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

    Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that support diverse future scene understanding tasks. …

  17. arXiv cs.AI TIER_1 English(EN) · Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang ·

    Beyond a single latent space: a dual-latent world model for long-horizon planning

    arXiv:2609.37644v1 Announce Type: new Abstract: Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discriminati…

  18. arXiv cs.AI TIER_1 English(EN) · Ke Fang, Yupu Yao, Lu Cheng ·

    ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning

    arXiv:2609.36333v1 Announce Type: cross Abstract: Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relev…

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning

    Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representa…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts

    World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-co…

  21. arXiv cs.CV TIER_1 English(EN) · Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis ·

    Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

    arXiv:2610.01942v1 Announce Type: new Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that su…

  22. arXiv cs.CV TIER_1 English(EN) · Ali Alrasheed, Basim Azam, Naveed Akhtar ·

    The Planning Limits of Latent World Models

    arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear…