PulseAugur
中
实时 23:56:11
English(EN) NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields

新的VLA模型增强机器人推理、安全性和效率 · 跟踪10个来源

研究人员正在为机器人技术开发先进的视觉-语言-动作(VLA)模型,重点是提高推理、安全性和效率。TOAST和TRUST等新方法旨在通过引入随机性和监控推理过程来增强策略学习和运行时安全性。其他进展包括用于城市导航的专用模型(UrbanVLA)、可泛化的奖励生成(Large Reward Models)以及节能的脉冲驱动架构。此外,还在努力改进VLA系统中的多智能体协调和长时任务规划。 AI

影响 这些进展推动了机器人控制和导航的边界,有可能带来更强大、更高效的自主系统。

排序理由 多篇研究论文介绍了用于视觉-语言-动作任务的新模型和方法。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 31 个来源。 我们如何撰写摘要 →

新的VLA模型增强机器人推理、安全性和效率 · 跟踪10个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了用于视觉-语言-动作任务的新模型和方法。
Source corroboration
31 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [31]

  1. arXiv cs.AI TIER_1 English(EN) · Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal ·

    当推理助力行动:在视觉-语言-行动策略中监控和引导思维链

    arXiv:2610.00601v1 Announce Type: cross Abstract: Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we de…

  2. arXiv cs.AI TIER_1 English(EN) · Yanru Wu, Weiduo Yuan, Esteban Martinez Licon, Ang Qi, Vitor Guizilini, Jiageng Mao, Yue Wang ·

    大型奖励模型:利用视觉语言模型实现可泛化的在线机器人奖励生成

    arXiv:2603.16065v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically mea…

  3. arXiv cs.AI TIER_1 English(EN) · Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang ·

    UrbanVLA:城市微出行领域的视觉-语言-动作模型

    arXiv:2510.23576v2 Announce Type: replace-cross Abstract: Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the …

  4. arXiv cs.AI TIER_1 English(EN) · Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae ·

    TOAST:用于自回归视觉-语言-动作模型的随机机器人动作标记化

    arXiv:2610.00899v1 Announce Type: cross Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation…

  5. arXiv cs.CL TIER_1 English(EN) · Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang ·

    Spike驱动的视觉-语言-动作模型

    arXiv:2609.39514v1 Announce Type: new Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy c…

  6. arXiv cs.AI TIER_1 English(EN) · Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li ·

    通过强化微调实现协同多智能体视觉-语言-动作模型

    arXiv:2609.36588v1 Announce Type: cross Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-gra…

  7. arXiv cs.AI TIER_1 English(EN) · Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid ·

    Gondola:机器人操作的地面视觉语言规划

    arXiv:2506.11261v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretabil…

  8. arXiv cs.AI TIER_1 English(EN) · Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar ·

    改变想法,改变行为:推理链作为视觉-语言-动作策略的控制面

    arXiv:2603.12717v2 Announce Type: replace-cross Abstract: Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are designed to reason in text before acting, generating a reasoning chain and decoding…

  9. arXiv cs.CL TIER_1 English(EN) · Haotian Deng, Wenbin Xing, Gang Xu, Tao He, Jinkai Zheng, Chun Li, Zheng Zhu, Ming Li ·

    视觉-语言模型在面对隐性风险时能否保持有用性?用于高效安全-有用性对齐的Intent-Privilege OPSD

    arXiv:2609.37837v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference dat…

  10. arXiv cs.LG TIER_1 English(EN) · Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng, Haibo Ding, Jun Huan ·

    StructRL:面向长时域视觉-语言-动作任务的在线结构化强化学习

    arXiv:2609.36352v1 Announce Type: cross Abstract: Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (…

  11. arXiv cs.LG TIER_1 English(EN) · Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang ·

    通过预测域不变潜在前瞻性来解耦视觉-语言-动作模型中的虚假相关性

    arXiv:2609.37165v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lo…

  12. arXiv cs.LG TIER_1 English(EN) · Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao ·

    面向高效视觉-语言-动作策略学习的对齐引导流Transformer

    arXiv:2609.34467v2 Announce Type: replace Abstract: Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due t…

  13. arXiv cs.AI TIER_1 English(EN) · Yunzhe Xu, Zhe Liu ·

    系统化多智能体视觉与语言导航:公式、基准和方法

    arXiv:2609.35965v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. W…

  14. arXiv cs.AI TIER_1 English(EN) · Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han ·

    面向视觉-语言-动作模型的线性表示假设

    arXiv:2609.30996v1 Announce Type: cross Abstract: The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun ext…

  15. arXiv cs.LG TIER_1 English(EN) · Ahad Jawaid, Yu Xiang ·

    NAC: 神经动作编码器用于视觉-语言-动作模型

    arXiv:2606.21372v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and down…

  16. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Simin Li ·

    通过强化微调实现协同多智能体视觉-语言-动作模型

    We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot …

  17. Hugging Face Daily Papers TIER_1 English(EN) ·

    StructRL:面向长时域视觉-语言-动作任务的在线结构化强化学习

    Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment…

  18. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Yujie Wu ·

    DS-VLA:一种受树突启发的视觉-语言-动作模型,用于鲁棒的动作控制

    Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a …

  19. arXiv cs.CV TIER_1 English(EN) · Guransh Singh ·

    AEGIS:用于知识保留的视觉-语言-动作微调的锚点强制梯度隔离

    arXiv:2604.16067v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continuous flow-matching action experts backpropagate concentrated, low-rank regression …

  20. arXiv cs.CV TIER_1 English(EN) · Yuliang Cai, Mohammad Rostami, Jesse Thomason ·

    Skeleton-and-Strategy Prompting:面向视觉语言模型的无训练否定理解

    arXiv:2610.01180v1 Announce Type: new Abstract: Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve…

  21. arXiv cs.CV TIER_1 English(EN) · Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu ·

    ATI-VLA:通过可操作的对齐和自适应注入实现以动作为中心的预测视觉-语言-动作模型

    arXiv:2610.01741v1 Announce Type: new Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct actio…

  22. arXiv cs.CV TIER_1 English(EN) · Shota Kobayashi, Koki Seno, Daichi Yashima, Komei Sugiura ·

    NarrativeFlow:基于流的视觉-语言-动作模型,使用机器人速度场

    arXiv:2610.00981v1 Announce Type: cross Abstract: We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This tas…

  23. arXiv cs.CV TIER_1 English(EN) · Hongyi Cai, Yi Herng Ong, Tingshiuan C. Wu, Chiew Hui Lim, Hanxia Li, Kehong Guo, Sze Yuan Cheong ·

    Devol-ONE: 一个自回归Transformer混合模型,用于统一视觉-语言-动作与潜在世界建模

    arXiv:2609.32193v2 Announce Type: replace Abstract: Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address th…

  24. arXiv cs.CV TIER_1 English(EN) · Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu ·

    DynamicVLA:用于动态对象操作的视觉-语言-动作模型

    arXiv:2601.22153v2 Announce Type: replace-cross Abstract: Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception-execution m…

  25. arXiv cs.CV TIER_1 English(EN) · Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie ·

    内存内联与可复用技能:面向视觉-语言-动作模型的以记忆为中心的框架

    arXiv:2609.39794v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring subst…

  26. arXiv cs.CV TIER_1 English(EN) · Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone ·

    具有语言记忆的视觉-语言-动作自主驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理自动驾驶代理

    arXiv:2609.38641v1 Announce Type: new Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable…

  27. arXiv cs.CV TIER_1 English(EN) · Jingqiu Wang, Yan Wang ·

    MotionWeave:为视觉-语言-动作策略学习以运动为中心的未来动力学

    arXiv:2609.39324v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant a…

  28. arXiv cs.CV TIER_1 English(EN) · Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi ·

    D$^2$-VLA:用于长时动态操作的双记忆双频视觉-语言-动作模型

    arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visua…

  29. arXiv cs.CV TIER_1 English(EN) · Songhua Yang, Ziyu Liu, Yuanwei Liu, Xuetao Li, Xuanye Fei, He Huang, Zheng Wang, Miao Li ·

    利用漏洞:机器人视觉-语言-动作模型通用对抗性攻击

    arXiv:2609.39178v1 Announce Type: cross Abstract: Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since …

  30. arXiv cs.CV TIER_1 English(EN) · Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin ·

    纠正“何处”,保留“如何”:通过指代引导实现视觉-语言-动作模型的组合泛化

    arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diver…

  31. arXiv cs.CV TIER_1 English(EN) · Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie ·

    NavHarness:Agentic视觉语言导航的自适应目标

    arXiv:2609.39915v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions doe…