Ant Group's Lingbo releases suite of embodied AI models, including world action and video generation
ByPulseAugur Editorial·[24 sources]·
Ant Group's Lingbo Technology has released a suite of new models aimed at advancing embodied AI and robotics. LingBot-VA 2.0 is presented as the first embodiment-native world action model, designed from the ground up for physical world interaction rather than adapting digital world models. This is complemented by LingBot-World 2.0, a real-time interactive world model capable of hour-long generation and incorporating AI agent mechanisms for dynamic interaction. Additionally, LingBot-Video, an MoE-based video generation model, is optimized for embodied AI tasks, outperforming existing models on robotics benchmarks. The company also launched LingBot-Depth 2.0 and LingBot-Vision for enhanced spatial perception and visual representation in robots.
AI
IMPACT
These releases advance the capabilities of embodied AI, potentially accelerating the development and deployment of more sophisticated robots capable of real-world interaction and task execution.
RANK_REASON
Multiple frontier-lab model releases (world action, world model, video generation, spatial perception) from Ant Group's Lingbo Technology.
Ant Group robotics unit Lingbo develops robot intelligence platform leveraging Ant's massive payment ecosystem data, taking an unconventional approach to embodied AI development.
<p>Ant Group's Robbyant has released the LingBot-VA 2.0 technical report — a Physical AI video-action foundation model built from scratch for embodiment rather than fine-tuned from a video generator. It predicts future states ahead of execution through Foresight Reasoning, re-gro…
<p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…
<p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…
<p>Ant Group's Robbyant open-sourced LingBot-Vision, a self-supervised ViT family for dense spatial perception. Masked boundary modeling makes image boundaries a native training signal. The 1B backbone matches or surpasses larger models, and initializes LingBot-Depth 2.0.</p> <p>…
Ant Group's LingBot-Depth 2.0 spatial perception model achieves 12 world-first benchmarks, with its 1.1B parameter LingBot-Vision foundation model surpassing Meta's 7B DINOv3.
Ant Group has unveiled LingBot-VA 2.0, a Physical AI video-action foundation model built natively for robotics. The model achieves 225 Hz asynchronous control and 93.6% average success rate on robot manipulation tasks. https://www. marktechpost.com/2026/07/11/an t-groups-robbyant…
Robbyant, Ant Group's AI unit, has released LingBot-World-Infinity - an open causal world model with an agentic harness. The 14B model uses a mixture of bidirectional and autoregressive attention to solve long-horizon drift. A 1.3B variant runs on a single GPU. https://www. markt…
Ant Group has open-sourced LingBot-Vision, a 1B vision model for robotics that treats object boundaries as a native training signal. It matches or surpasses models 7x larger on spatial perception tasks. Available on Hugging Face under Apache-2.0. https://www. marktechpost.com/202…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up47qv/ant_group_released_lingbotvision_dinofamily/"> <img alt="Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer …