English(EN)PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving
新的VLA框架推动自动驾驶感知和动作规划 · 跟踪9个来源
作者PulseAugur 编辑部·[10 个来源]·
多篇研究论文介绍了用于自动驾驶的新型框架,这些框架集成了视觉、语言和动作(VLA)能力。MATS提出了一种多模态、多任务学习方法,具有自适应融合和特定任务的专家,用于3D感知。HyWorldVLA结合了像素级监督和基于潜在变量的世界模型以实现鲁棒驾驶,而PerceptDrive则利用冻结的感知模型和自适应专家路由。ForgeDrive使用统一的扩散框架和交叉条件进行视觉-动作生成,MindDrive则采用在线强化学习和大型语言模型进行决策。此外,反事实视觉动作分析(CVAA)提供了一种通过分析VLA模型对物体移除的响应来解释其模型的方法。
AI
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from m…
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan trans…
arXiv cs.AI
TIER_1English(EN)·Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer·
arXiv:2607.16938v1 Announce Type: cross Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet t…
arXiv:2607.15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the …
End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems…
arXiv cs.CV
TIER_1English(EN)·Xuchang Zhong, He Zheng, Chenxu Zhao, Tianxiong Lv, Hangqi Fan, Bohua Wang, Yushan Liu, Li Gao, Zhihao Liao, Leigang Luo, Congyang Zhao, Yang Cai·
arXiv:2606.31226v2 Announce Type: replace Abstract: World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visu…
arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex f…
arXiv:2607.20175v1 Announce Type: new Abstract: Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. W…
arXiv:2512.13636v4 Announce Type: replace Abstract: Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement Learning …