PulseAugur
实时 05:20:32
English(EN) OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving

OmniDrive-R1 通过强化驱动的视觉基础增强自动驾驶 VLMs

研究人员推出 OmniDrive-R1,一个用于自动驾驶的新型框架,它使用交错多模态思维链 (iMCoT) 机制整合感知和推理。该方法通过采用强化驱动的视觉基础能力,解决了视觉语言模型中常见的对象幻觉问题。该系统利用独特的无标注训练流程和 Clip-GRPO 算法,该算法在不需要密集定位标签的情况下生成基础奖励。实验表明,与基线模型相比,OmniDrive-R1 显著提高了推理分数和准确性。 AI

影响 引入了一种新颖的方法来提高安全关键型自动驾驶应用中 VLM 的可靠性。

排序理由 这是一篇详细介绍自动驾驶新模型和方法的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

OmniDrive-R1 通过强化驱动的视觉基础增强自动驾驶 VLMs

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Zhenguo Zhang, Haohan Zheng, Yishen Wang, Le Xu, Tianchen Deng, Xuefeng Chen, Qu Chen, Bo Zhang, Wuxiong Huang ·

    OmniDrive-R1:用于可信视觉-语言自动驾驶的强化驱动交错多模态思维链

    arXiv:2512.14044v3 Announce Type: replace-cross Abstract: The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability failures, most notably object hallucination. This failure stems from their relia…