研究人员正在探索具身智能的进展,重点关注连接机器人感知和决策的“世界模型”。论文讨论了从“合理”到“可操作”对这些模型进行分类的框架,强调它们在改善机器人行为和任务执行方面的作用。FluxVLA Engine等平台和Pelican-Sim 1.0等模拟器正在开发中,以简化这些复杂系统的工程和部署,解决数据集成、训练和实际应用中的挑战。
AI
arXiv:2609.19659v1 Announce Type: cross Abstract: Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samp…
arXiv cs.LG
TIER_1English(EN)·Haoqiang Kang, Yiming Zhang, Yiyang Guo, Chuying Li, Jianzhi Shen, Tianruo Rose Xu, Xiaokang Ye, Lianhui Qin·
arXiv:2609.19801v1 Announce Type: new Abstract: Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, …
Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, or money needed for later work. Learning to plan…
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise lo…
arXiv:2609.16697v1 Announce Type: cross Abstract: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress…
arXiv:2609.17210v1 Announce Type: cross Abstract: Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems rema…
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions to follow instructions or locate objects. Training-based methods face generalization challenges, while training-free methods exploit multimodal large language m…
Pelican-Sim 1.0 is a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions, using unified action representations, action-visual injection, sparse mixture-of-experts, and efficient rollout generation to impr…
arXiv:2609.19554v1 Announce Type: cross Abstract: Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to …
Songyan Dynamics' Scalabot brand released HERON-CRA, pairing 4D Context Expert memory, a plug-in RL Engine, and cross-embodiment pretraining, with sock-folding success rising from 38.5% to 97.8%.