English(EN)Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
新基准和方法增强多模态AI推理和可信度 · 跟踪4个来源
作者PulseAugur 编辑部·[4 个来源]·
研究人员正在开发新方法来提高多模态大语言模型(MLLMs)的可靠性和可信度。一种方法VERDICT利用多个验证器之间的分歧来识别推理步骤中的错误,而无需额外的训练数据。另一种方法AD2-Bench引入了一个分层诊断框架,以查明证据获取中的失败,区分空间歧义和语义不确定性。此外,StructReward提供了一个高效的框架,通过提供结构化的、步骤级别的奖励来实现多模态推理的自我纠正,降低了强化学习的计算开销。最后,MMArch提供了一个专门针对建筑和土木工程领域多模态推理的基准,突显了当前MLLMs在应用原理和结合证据方面与人类专家表现之间存在的显著差距。
AI
arXiv:2608.10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelle…
arXiv:2608.10954v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models of…
arXiv:2608.08326v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answ…
arXiv cs.AI
TIER_1English(EN)·Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren·
arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distr…