PulseAugur
EN
LIVE 10:45:58

New simulators and frameworks advance embodied AI research and deployment

Researchers are developing advanced simulators and frameworks to enhance embodied AI research and deployment. SPEAR, a Python library, offers high-speed, photorealistic rendering and extensive programmability for Unreal Engine applications. Another paper introduces "Embodied Operators" as reusable modules for embodied intelligence systems, proposing a benchmark framework for their evaluation. Additionally, a new approach called Agent Architecture Search (AAS) automates the design of embodied agent architectures, moving beyond manual construction. Efforts are also underway to create unified frameworks for data collection, inference, and deployment on real robots, such as EVA-Client and Embodied.cpp, enabling more efficient and portable execution of embodied AI models. AI

IMPACT Advances in simulators and deployment frameworks are crucial for accelerating the development and real-world application of embodied AI systems.

RANK_REASON Multiple research papers introducing new simulators, frameworks, and methodologies for embodied AI.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 38 sources. How we write summaries →

New simulators and frameworks advance embodied AI research and deployment

COVERAGE [38]

  1. arXiv cs.AI TIER_1 English(EN) · Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He ·

    UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

    arXiv:2607.13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks…

  2. arXiv cs.AI TIER_1 English(EN) · Keji He ·

    UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

    Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often ne…

  3. arXiv cs.AI TIER_1 English(EN) · Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huapi… ·

    Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

    arXiv:2607.11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, an…

  4. arXiv cs.AI TIER_1 English(EN) · Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, Xuelong Li ·

    From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    arXiv:2607.11689v1 Announce Type: cross Abstract: Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) a…

  5. arXiv cs.AI TIER_1 English(EN) · Xuelong Li ·

    From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect can…

  6. arXiv cs.AI TIER_1 English(EN) · Jason Li ·

    Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

    Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods t…

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

    Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods t…

  8. arXiv cs.AI TIER_1 English(EN) · Mike Roberts, Renhan Wang, Rushikesh Zawar, Rachith Dey-Prakash, Quentin Leboutet, Stephan R. Richter, Matthias M\"uller, German Ros, Rui Tang, Stefan Leutenegger, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Vladlen Koltun ·

    SPEAR: A Simulator for Photorealistic Embodied AI Research

    arXiv:2607.06701v1 Announce Type: cross Abstract: Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We a…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    SPEAR: A Simulator for Photorealistic Embodied AI Research

    Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A S…

  10. arXiv cs.AI TIER_1 English(EN) · Junwu Xiong, Jiaxuan Gao, Wei Chai, Renxing Chen, Yuzhen Li, Yu Guo, Yucheng Guo, Mingxi Luo, Wenyang Ma, Yiyun Mou, Yifei Zhang, Chen Zhou, Yongjian Guo ·

    Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

    arXiv:2607.03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representati…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    Automating the Design of Embodied Agent Architectures

    Automated agent architecture search demonstrates potential for improving embodied agent performance while revealing challenges related to optimization signals, local optima, and credit assignment in simulation-based training.

  12. 量子位 (QbitAI) TIER_1 中文(ZH) · 林, 方舟 ·

    World Models Are Here: A Benchmark for Causal Technology! Embodied Brains Are Really Getting Brains!

  13. arXiv cs.AI TIER_1 English(EN) · Chengyang Li, Yikun Wang, Jiahui He, Yujie Wan, Shuai Wang, Yuan Wu, Yik-Chung Wu, Chengzhong Xu, Huseyin Arslan ·

    Memory-Native Non-Terrestrial Networks for Embodied Intelligence

    arXiv:2607.00029v1 Announce Type: cross Abstract: Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in wilderness to leverage cloud resources or report critical information to remote centers. However, the synergy is non…

  14. arXiv cs.AI TIER_1 English(EN) · Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, Xudong Xu ·

    EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    arXiv:2604.01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric sim…

  15. arXiv cs.AI TIER_1 English(EN) · Jinwoo Jang, Daniel J. Rho, Sihyung Yoon, Hyunsuk Cho, Honguk Woo ·

    Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

    arXiv:2607.00457v1 Announce Type: new Abstract: Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in applying Mixture of Experts (MoE) to this setting: routing lacks an explicit noti…

  16. Hugging Face Daily Papers TIER_1 English(EN) ·

    Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

    Embodied.cpp is a portable C++ runtime that enables efficient deployment of vision-language-action and world-action models across heterogeneous edge devices through modular execution layers and optimized inference.

  17. Hugging Face Daily Papers TIER_1 English(EN) ·

    EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

    EVA-Client is an open-source framework that unifies real-robot policy deployment, data collection, and evaluation through a component-decoupled architecture with inspectable execution workflows.

  18. 量子位 (QbitAI) TIER_1 中文(ZH) · henry ·

    Embodied Intelligence Skill Moment! NVIDIA Open-Sources Robot Skill Library, Jim Fan: The Paradigm Has Shifted

    全新的持续学习范式

  19. arXiv cs.AI TIER_1 English(EN) · Honguk Woo ·

    Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

    Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in applying Mixture of Experts (MoE) to this setting: routing lacks an explicit notion of scale, preventing targeted updates at spec…

  20. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    Vesta: A Generalist Embodied Reasoning Model

    Vesta is a unified embodied generalist model that integrates localization, spatial reasoning, navigation, and long-horizon planning into a single foundation model, outperforming specialized models in both benchmark tests and real-world robotic applications.

  21. arXiv cs.CV TIER_1 English(EN) · Xinjie Wang, Liu Liu, Taojun Ding, Andrew Choi, Chaodong Huang, Mengao Zhao, Ziang Li, Jackson Jiang, Chunlei Yu, Shengxiang Liu, Wei Xu, Zhizhong Su ·

    EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

    arXiv:2607.07459v1 Announce Type: cross Abstract: We present EmbodiedGen V2, a generative 3D world engine for building executable sim-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapidly, yet assembling such assets into policy-ready tas…

  22. arXiv cs.CV TIER_1 English(EN) · Zhizhong Su ·

    EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

    We present EmbodiedGen V2, a generative 3D world engine for building executable sim-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapidly, yet assembling such assets into policy-ready task environments remains largely manual, limiting sc…

  23. arXiv cs.CV TIER_1 English(EN) · Heqing Yang, Yang Yi, Liyao Wang, Linqing Zhong, Donglin Yang, Ruipu Wu, Zitong Bai, Fengjiao Chen, Manyuan Zhang, Linjiang Huang, Si Liu ·

    EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

    arXiv:2607.02646v1 Announce Type: cross Abstract: We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and the physical hardware, EVA-Client unifies the rea…

  24. arXiv cs.CV TIER_1 English(EN) · Yifan Li, Lichi Li, Anh Dao, Xinyu Zhou, Wenjun Huang, Tianyi Ma, Yicheng Qiao, Zheda Mai, Daeun Lee, Zichen Chen, Pan Wang, Lehan Yang, Tianlong Wang, Zhen Tan, Sheng Li, Mohit Bansal, Yang Ni, Yu Kong ·

    IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

    arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reasoning. Existing embodied benchmarks largely focus on passive, static household e…

  25. arXiv cs.CV TIER_1 English(EN) · Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang ·

    Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

    arXiv:2607.02501v1 Announce Type: cross Abstract: Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especi…

  26. arXiv cs.CV TIER_1 English(EN) · Shuai Wang ·

    Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

    Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing infer…

  27. arXiv cs.CV TIER_1 English(EN) · Liqiang Nie ·

    Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

    Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to the RoboSpatial Challenge at the Embodied Reasoning in Action Workshop, CVPR 2026, built on RoboBrain…

  28. arXiv cs.CV TIER_1 English(EN) · Longyu Chen, Heng Li, Wei Yang, Manqi Zhao, Dongsheng Jiang ·

    ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

    arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluation of Vision-Language-Action systems. However, their reliability as evaluation …

  29. arXiv cs.CV TIER_1 English(EN) · Xinqing Li, Xin He, Le Zhang, Min Wu, Xiaoli Li, Yun Liu ·

    A Comprehensive Survey on World Models for Embodied AI

    arXiv:2510.16732v3 Announce Type: replace Abstract: Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to…

  30. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Building a 'Data Factory' for Robots, How Xiaomi Robotics-U0 Solves the Toughest Problem in Embodied Intelligence?

    <section><section><section><section><section></section><section><section><section><section></section></section></section><section><span>世界各地,有不少人正在给机器人当“幼教”。</span></section><p style="text-align: justify; margin-bottom: 24px; line-height: 1.75em;"><span><span style="letter-spacin…

  31. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Breaking Through the Generalization Bottleneck of Embodied Intelligence! Ant Lingbo Open-Sources Embodied Foundation Model LingBot-VLA 2.0, Supporting Over 20 Robot Configurations

    <p>7&nbsp;月&nbsp;8&nbsp;日,蚂蚁灵波科技宣布升级并开源新一代具身基座模型&nbsp;LingBot-VLA 2.0。作为今年&nbsp;1&nbsp;月开源版本&nbsp;LingBot-VLA 1.0&nbsp;的全面升级,LingBot-VLA 2.0在预训练阶段融入6万小时高质量真实物理数据,覆盖&nbsp;17&nbsp;个主流机器人品牌的&nbsp;20&nbsp;种机器人构型,并扩展对头部、腰部、末端执行器及移动底盘等自由度的支持。在构型泛化、自由度支持和落地效率等方面实现显著提升。&nbsp;</p><p>当前具身智…

  32. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Amio Robot Liu Fang: Embodied intelligence is not a large model, let alone intelligent driving

    <p>过去两年,具身智能常被放进两个熟悉的叙事里理解:它要么是“大模型的下一站”,要么是“自动驾驶向工厂的迁移”。前者强调模型、数据和规模效应,后者强调感知、预测、规划与端到端闭环。两种说法都言之成理,然而在阿米奥机器人创始人刘方看来,它们仍未能回答那个更根本的问题——我们究竟需要具身智能解决什么问题。</p><p>刘方坚信:具身不是大模型,不是自动驾驶,是新业态。自动驾驶数字化的是驾驶能力;大模型数字化的是知识;具身智能真正要数字化的,是劳动能力。</p><p>这并不是一句把机器人包装得更宏大的口号。相反,它把问题拉回了最务实的工业现场:客户购买的不是…

  33. SCMP — Tech TIER_1 English(EN) · Minxiao Chang ·

    The next frontier of AI: how ‘world models’ are simulating reality and virtual spaces

    Artificial intelligence has spent years learning to generate words, images and videos. Now, the industry’s attention is shifting towards “world models” – AI systems designed to simulate how the physical or digital world changes over time. While the term was originally associated …

  34. Forbes — Innovation TIER_1 English(EN) · Michael Ashley, Contributor ·

    Is Embodied AI The Next Great Computing Revolution?

    Embodied AI offers the next computing revolution. Here's why executives should look beyond LLMs and prepare for robots that learn by interacting with the physical world.

  35. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Xiaomi Open-Sources Robotics-U0: A 38B-Parameter Embodied Generative Model That Unifies Four Robot Tasks

    Xiaomi releases Robotics-U0 with 38 billion parameters, the first embodied generative model handling scene generation, embodied transfer, video generation, and text-to-image in a single unified architecture.

  36. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Every 48 Hours, a New Embodied AI Model Is Born: From BAAI World Model to Alibaba Qwen-Robot

    June 2026 saw 13 new embodied AI models and world models released, tracking the shift from hardware benchmarks to software intelligence competition in embodied AI.

  37. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Embodied AI's Spatial Vision Achilles' Heel Finally Solved: CMG Lab's IROS 2026 Breakthrough

    China Merchants Group's LionRock AI Lab identifies and solves VLA shortcut learning with Hybrid Dynamic Data Collection, accepted at IROS 2026.

  38. AI Business TIER_1 English(EN) · Scarlett Evans ·

    Chinese Tech Vendors Converge on Humanoid Robotics and Embodied AI

    China is racing to stake a claim in the fast-growing sector.