蚂蚁集团凌波科技发布了一系列旨在推进具身AI和机器人技术的新模型。LingBot-VA 2.0被呈现为首个具身原生世界动作模型,从根本上为物理世界交互而设计,而非改编数字世界模型。与之相辅的是LingBot-World 2.0,一个能够进行长达一小时生成的实时交互式世界模型,并整合了AI代理机制以实现动态交互。此外,LingBot-Video,一个基于MoE的视频生成模型,针对具身AI任务进行了优化,在机器人基准测试中表现优于现有模型。该公司还推出了LingBot-Depth 2.0和LingBot-Vision,以增强机器人的空间感知和视觉表示能力。
AI
Ant Group robotics unit Lingbo develops robot intelligence platform leveraging Ant's massive payment ecosystem data, taking an unconventional approach to embodied AI development.
<p>Ant Group's Robbyant has released the LingBot-VA 2.0 technical report — a Physical AI video-action foundation model built from scratch for embodiment rather than fine-tuned from a video generator. It predicts future states ahead of execution through Foresight Reasoning, re-gro…
<p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…
<p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…
<p>Ant Group's Robbyant open-sourced LingBot-Vision, a self-supervised ViT family for dense spatial perception. Masked boundary modeling makes image boundaries a native training signal. The 1B backbone matches or surpasses larger models, and initializes LingBot-Depth 2.0.</p> <p>…
Ant Group's LingBot-Depth 2.0 spatial perception model achieves 12 world-first benchmarks, with its 1.1B parameter LingBot-Vision foundation model surpassing Meta's 7B DINOv3.
Ant Group has unveiled LingBot-VA 2.0, a Physical AI video-action foundation model built natively for robotics. The model achieves 225 Hz asynchronous control and 93.6% average success rate on robot manipulation tasks. https://www. marktechpost.com/2026/07/11/an t-groups-robbyant…
Robbyant, Ant Group's AI unit, has released LingBot-World-Infinity - an open causal world model with an agentic harness. The 14B model uses a mixture of bidirectional and autoregressive attention to solve long-horizon drift. A 1.3B variant runs on a single GPU. https://www. markt…
Ant Group has open-sourced LingBot-Vision, a 1B vision model for robotics that treats object boundaries as a native training signal. It matches or surpasses models 7x larger on spatial perception tasks. Available on Hugging Face under Apache-2.0. https://www. marktechpost.com/202…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up47qv/ant_group_released_lingbotvision_dinofamily/"> <img alt="Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer …