PulseAugur
中
实时 15:56:59
English(EN) Native Video-Action Pretraining for Generalizable Robot Control

LingBot-VA 2.0:面向通用机器人控制的新基础模型

研究人员推出 LingBot-VA 2.0,这是一款专为物理环境中的机器人控制设计的新型视频-动作基础模型。与从数字内容生成改编而来的模型不同,LingBot-VA 2.0 采用语义视觉-动作分词器,以提高指令遵循和动作精度。它还利用因果预训练范式来防止灾难性遗忘,并采用稀疏专家混合(MoE)骨干网络以实现高效高频推理。该模型在现实世界部署中验证了其实时闭环控制能力,在复杂操作任务中展示了强大的少样本泛化能力。 AI

影响 该模型针对物理环境的专业设计及其展示的泛化能力,有望加速开发更强大、更具适应性的机器人。

排序理由 该集群包含一篇详细介绍新模型及其能力的学术论文。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

LingBot-VA 2.0:面向通用机器人控制的新基础模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇详细介绍新模型及其能力的学术论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
90 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.CV TIER_1 English(EN) · Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang, Yiming Luo, Shuaiting Li, Ruilin Wang, Junke Wang, Jiahao Shao, Gangwei Xu, Jiaming Zhou, Yishu Shen, Yudong Jin, Fangyi Xu, Shuailei Ma, Jiaqi Liao, Guanxing Lu, Zifan Shi, Yongkun Wen, Yujie Zhao, Weixuan … ·

    面向通用机器人控制的原生视频动作预训练

    arXiv:2607.08639v1 Announce Type: cross Abstract: The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments.…

  2. arXiv cs.CV TIER_1 English(EN) · Yinghao Xu ·

    面向通用机器人控制的原生视频动作预训练

    The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a …