PulseAugur
EN
LIVE 21:34:22
中文(ZH) 机器人视觉迎来新突破!蚂蚁灵波空间感知模型LingBot-Depth 2.0正式发布

Ant Group's Lingbo releases suite of embodied AI models, including world action and video generation

Ant Group's Lingbo Technology has released a suite of new models aimed at advancing embodied AI and robotics. LingBot-VA 2.0 is presented as the first embodiment-native world action model, designed from the ground up for physical world interaction rather than adapting digital world models. This is complemented by LingBot-World 2.0, a real-time interactive world model capable of hour-long generation and incorporating AI agent mechanisms for dynamic interaction. Additionally, LingBot-Video, an MoE-based video generation model, is optimized for embodied AI tasks, outperforming existing models on robotics benchmarks. The company also launched LingBot-Depth 2.0 and LingBot-Vision for enhanced spatial perception and visual representation in robots. AI

IMPACT These releases advance the capabilities of embodied AI, potentially accelerating the development and deployment of more sophisticated robots capable of real-world interaction and task execution.

RANK_REASON Multiple frontier-lab model releases (world action, world model, video generation, spatial perception) from Ant Group's Lingbo Technology.

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 24 sources. How we write summaries →

Ant Group's Lingbo releases suite of embodied AI models, including world action and video generation

COVERAGE [24]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Industry's first embodied native world action model is here! Ant Lingbo releases LingBot-VA 2.0

  2. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    World Models First Achieve 'Hour-Level' Generation! Ant Lingbo Open-Sources LingBot-World 2.0, Supporting AI-Native Multi-Person Interaction

  3. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Ant Lingbo open-sources LingBot-Video, the world's first embodied video foundation model is here!

  4. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Robotic Vision Welcomes New Breakthrough! Ant Lingbo Spatial Perception Model LingBot-Depth 2.0 Officially Released

  5. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Industry's first embodied native world action model is here! Ant Lingbo releases LingBot-VA 2.0

    <p>7 月 10 日,蚂蚁灵波发布业界首个具身原生世界动作模型 LingBot-VA 2.0。该模型的发布,标志着机器人基础模型正式从“基于数字世界模型构建”到“面向物理世界原生设计”的关键转变。它代表了具身智能发展的一种关键路线选择:机器人“大脑”不再依托数字世界模型能力的“嫁接”,而是从动态建模、因果预测、实时执行等与环境交互的原始需求出发,进行原生设计。</p><p>&nbsp;</p><p>得益于具身原生架构,LingBot-VA 2.0在真机测试中表现出了出色的执行速度和泛化能力。以下面这个视频为例,在不依赖任何外部拍摄设备的情况下,机器人就…

  6. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Ant Group Releases LingBot-VA 2.0

    36氪获悉,7月10日,蚂蚁灵波发布业界首个具身原生世界动作模型LingBot-VA 2.0。它代表了具身智能发展的一种关键路线选择:机器人“大脑”不再依托数字世界模型能力的“嫁接”,而是从动态建模、因果预测、实时执行等与环境交互的原始需求出发,进行原生设计。

  7. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    China's Robot Brain Ignites Overseas, Ant Lingbo Full Stack 2.0 Model Defines Embodied Intelligence "New Brain"

    <p>7月9日,中国AI创新力量再次引发全球关注。</p><p>中国具身智能公司——蚂蚁灵波科技升级全栈大脑2.0,一口气发布了涵盖视觉、动作、操作以及物理预测等方向的具身智能大模型,在海外开发者社区掀起热潮,发布即登上海外社交媒体热榜。</p><p style="text-align: center;"><img src="https://static.leiphone.com/uploads/new/images/20260709/6a4f3b437431c.png?imageView2/2/w/740" /></p><p style="text-a…

  8. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Ant Lingbo open sources LingBot-Video, the world's first embodied video foundation model is here!

    <p>&nbsp;</p><p>7 月 9 日,蚂蚁灵波开源&nbsp;LingBot-Video,这是全球首个基于 Mixture-of-Experts(MoE)架构、面向具身智能的开源视频生成基础模型。该模型围绕机器人和具身智能的核心需求重新设计视频预训练范式,在推理效率、物理合理性、动作理解和任务完成度等方面取得系统性提升,为视频基础模型从数字内容创作走向具身智能提供了新的开源底座。</p><p>&nbsp;</p><p>在北京大学联合字节跳动发布的基准&nbsp;RBench&nbsp;上,LingBot-Video 的总分是 0.620,超越了…

  9. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Ant Lingbo Technology Open Sources Real-time Interactive World Model LingBot-World 2.0

    36氪获悉,蚂蚁灵波科技宣布开源新一代实时交互世界模型LingBot-World 2.0。该模型全面升级世界预测与交互能力,支持小时级实时生成、720p/60fps高清实时输出以及更丰富的交互动作与事件,并在业界首次将Agent机制引入世界模型。LingBot-World 2.0不仅可广泛应用于游戏内容生成、影视预演、虚拟仿真、数字孪生等场景,同时支持机器人与具身智能训练场景。

  10. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Official Announcement! Ant Lingbo World Model 2.0 is here, hour-level generation + Agent-driven defines a new paradigm for AI interaction

    <p>7&nbsp;月&nbsp;9&nbsp;日,蚂蚁灵波科技开源新一代实时交互世界模型&nbsp;LingBot-World 2.0(又称&nbsp;LingBot-World-Infinity)。该模型全面升级世界预测与交互能力,支持小时级实时生成、720p/60fps&nbsp;高清实时输出以及更丰富的交互动作与事件,并在业界首次将&nbsp;Agent&nbsp;机制引入世界模型,让生成的世界从“可观看、可操控”,进一步走向“可持续互动、可动态变化”。&nbsp;</p><p style="text-align: center;"><img s…

  11. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Ant Lingbo open sources LingBot-Video, the world's first embodied video foundation model is here!

    <p>7&nbsp;月&nbsp;9&nbsp;日,蚂蚁灵波开源&nbsp;LingBot-Video,这是全球首个基于&nbsp;Mixture-of-Experts(MoE)架构、面向具身智能的开源视频生成基础模型。该模型围绕机器人和具身智能的核心需求重新设计视频预训练范式,在推理效率、物理合理性、动作理解和任务完成度等方面取得系统性提升,为视频基础模型从数字内容创作走向具身智能提供了新的开源底座。</p><p>在北京大学联合字节跳动发布的基准&nbsp;RBench&nbsp;上,LingBot-Video&nbsp;的总分是&nbsp;0.620…

  12. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Ant Lingbo Embodied Foundation Model LingBot-VLA 2.0 Open Source

    7月8日,蚂蚁灵波科技宣布升级并开源新一代具身基座模型LingBot-VLA 2.0。作为今年1月开源版本LingBot-VLA 1.0的全面升级,LingBot-VLA 2.0在预训练阶段融入6万小时高质量真实物理数据,覆盖17个主流机器人品牌的20种机器人构型,并扩展对头部、腰部、末端执行器及移动底盘等自由度的支持。在构型泛化、自由度支持和落地效率等方面实现显著提升。

  13. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Robotic Vision Welcomes New Breakthrough! Ant Lingbo Spatial Perception Model LingBot-Depth 2.0 Officially Released

    <p>7月7日,蚂蚁集团旗下具身智能公司灵波科技发布空间感知模型 LingBot-Depth 2.0。该模型基于 1.5 亿规模数据进行训练,在边缘清晰度、细小物体识别、远距离深度估计以及复杂场景鲁棒性等方面实现全面升级。</p><p>LingBot-Depth是灵波自研的空间感知模型,相当于机器人在物理世界的眼睛,1.0版解决了机器人看清透明、反光等复杂场景的空间感知难题。相比于LingBot-Depth 1.0,LingBot-Depth 2.0的训练数据从300万扩充至1.5亿规模,性能全面升级:在深度补全基准的16项测评中获得12 项第一;在最难…

  14. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Ant Group releases spatial perception model LingBot-Depth 2.0

    7月7日,蚂蚁集团旗下具身智能公司灵波科技发布空间感知模型LingBot-Depth 2.0。该模型基于1.5亿规模数据进行训练,在边缘清晰度、细小物体识别、远距离深度估计以及复杂场景鲁棒性等方面实现全面升级。此外,灵波还同步推出了LingBot-Depth 2.0的视觉基座模型——LingBot-Vision,构建起机器人从“看懂”到“看准”的能力链路,旨在应对机器人视觉在空间感知、精细识别和复杂环境适应等方面的核心挑战。

  15. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Ant Group Robotics Subsidiary Lingbo Takes a New Approach to Building Robot Brains

    Ant Group robotics unit Lingbo develops robot intelligence platform leveraging Ant's massive payment ecosystem data, taking an unconventional approach to embodied AI development.

  16. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Ant Group’s Robbyant Unveils LingBot-VA 2.0: A Causal Video-Action Model Built Natively for Physical AI

    <p>Ant Group's Robbyant has released the LingBot-VA 2.0 technical report — a Physical AI video-action foundation model built from scratch for embodiment rather than fine-tuned from a video generator. It predicts future states ahead of execution through Foresight Reasoning, re-gro…

  17. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

    <p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…

  18. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

    <p>Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations an…

  19. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Ant Group’s Robbyant Open-Sources LingBot-Vision: A 1B Boundary-Centric Vision Foundation Model for Dense Spatial Perception

    <p>Ant Group's Robbyant open-sourced LingBot-Vision, a self-supervised ViT family for dense spatial perception. Masked boundary modeling makes image boundaries a native training signal. The 1B backbone matches or surpasses larger models, and initializes LingBot-Depth 2.0.</p> <p>…

  20. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Ant Group's LingBot-Vision Claims 12 World Firsts: 1.1B Parameter Model Beats 7B DINOv3

    Ant Group's LingBot-Depth 2.0 spatial perception model achieves 12 world-first benchmarks, with its 1.1B parameter LingBot-Vision foundation model surpassing Meta's 7B DINOv3.

  21. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Ant Group has unveiled LingBot-VA 2.0, a Physical AI video-action foundation model built natively for robotics. The model achieves 225 Hz asynchronous control a

    Ant Group has unveiled LingBot-VA 2.0, a Physical AI video-action foundation model built natively for robotics. The model achieves 225 Hz asynchronous control and 93.6% average success rate on robot manipulation tasks. https://www. marktechpost.com/2026/07/11/an t-groups-robbyant…

  22. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Robbyant, Ant Group's AI unit, has released LingBot-World-Infinity - an open causal world model with an agentic harness. The 14B model uses a mixture of bidirec

    Robbyant, Ant Group's AI unit, has released LingBot-World-Infinity - an open causal world model with an agentic harness. The 14B model uses a mixture of bidirectional and autoregressive attention to solve long-horizon drift. A 1.3B variant runs on a single GPU. https://www. markt…

  23. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Ant Group has open-sourced LingBot-Vision, a 1B vision model for robotics that treats object boundaries as a native training signal. It matches or surpasses mod

    Ant Group has open-sourced LingBot-Vision, a 1B vision model for robotics that treats object boundaries as a native training signal. It matches or surpasses models 7x larger on spatial perception tasks. Available on Hugging Face under Apache-2.0. https://www. marktechpost.com/202…

  24. r/LocalLLaMA TIER_1 English(EN) · /u/Simple_Response8041 ·

    Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up47qv/ant_group_released_lingbotvision_dinofamily/"> <img alt="Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer …