PulseAugur
实时 01:19:25
English(EN) PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

新的VLA框架推动自动驾驶感知和动作规划 · 跟踪9个来源

多篇研究论文介绍了用于自动驾驶的新型框架,这些框架集成了视觉、语言和动作(VLA)能力。MATS提出了一种多模态、多任务学习方法,具有自适应融合和特定任务的专家,用于3D感知。HyWorldVLA结合了像素级监督和基于潜在变量的世界模型以实现鲁棒驾驶,而PerceptDrive则利用冻结的感知模型和自适应专家路由。ForgeDrive使用统一的扩散框架和交叉条件进行视觉-动作生成,MindDrive则采用在线强化学习和大型语言模型进行决策。此外,反事实视觉动作分析(CVAA)提供了一种通过分析VLA模型对物体移除的响应来解释其模型的方法。 AI

影响 这些多样化的VLA框架通过改进感知、世界建模和动作规划,推动了自动驾驶的边界,有望实现更安全、更鲁棒的自动驾驶系统。

排序理由 多篇研究论文介绍了用于自动驾驶的新型框架。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 10 个来源。 我们如何撰写摘要 →

新的VLA框架推动自动驾驶感知和动作规划 · 跟踪9个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了用于自动驾驶的新型框架。
Source corroboration
10 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
49 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [10]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MATS:一种用于自动驾驶中3D感知的多模态多任务学习新框架

    Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from m…

  2. arXiv cs.AI TIER_1 English(EN) · Quanfu Yu, Xian Wu, Hao Xu, Liulong Ma ·

    HyWorldVLA:一个具有混合世界模型的自动驾驶视觉-语言-动作模型

    arXiv:2607.20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    PerceptDrive:用于端到端自动驾驶的具有自适应专家路由的感知先验世界-动作建模

    Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan trans…

  4. arXiv cs.AI TIER_1 English(EN) · Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer ·

    它们看到了什么?通过视觉-语言-动作模型解读复杂道路场景,以实现安全可信的自动驾驶汽车学习

    arXiv:2607.16938v1 Announce Type: cross Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet t…

  5. arXiv cs.AI TIER_1 English(EN) · Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada ·

    以 5 赫兹思考,以 20 赫兹行动:闭环驾驶的异步快慢视觉-语言-动作推理

    arXiv:2607.15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the …

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    它们看到了什么?通过视觉-语言-动作模型解读复杂道路场景,以实现安全可信的自动驾驶汽车学习

    End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems…

  7. arXiv cs.CV TIER_1 English(EN) · Xuchang Zhong, He Zheng, Chenxu Zhao, Tianxiong Lv, Hangqi Fan, Bohua Wang, Yushan Liu, Li Gao, Zhihao Liao, Leigang Luo, Congyang Zhao, Yang Cai ·

    ForgeDrive:自动驾驶中统一视觉-动作生成的双向交叉条件

    arXiv:2606.31226v2 Announce Type: replace Abstract: World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visu…

  8. arXiv cs.CV TIER_1 English(EN) · Junchen Huo, Wanming Hao, Song Wang, Enqing Chen, Shouyi Yang, Guanghui Wang ·

    MATS:一种用于自动驾驶中3D感知的多模态多任务学习新框架

    arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex f…

  9. arXiv cs.CV TIER_1 English(EN) · Yushan Liu, Tianxiong Lv, Bohua Wang, Hangqi Fan, Chenxu Zhao, He Zheng, Xuchang Zhong, Yifan Xie, Congyang Zhao, Zhihao Liao, Leigang Luo, Yang Cai, Xiao-Ping Zhang, Wenbo Ding ·

    PerceptDrive:用于端到端自动驾驶的感知先验世界-动作建模与自适应专家路由

    arXiv:2607.20175v1 Announce Type: new Abstract: Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. W…

  10. arXiv cs.CV TIER_1 English(EN) · Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Dingkang Liang, Xiang Bai ·

    MindDrive:一种通过在线强化学习实现自动驾驶的视觉-语言-动作模型

    arXiv:2512.13636v4 Announce Type: replace Abstract: Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement Learning …