PulseAugur
实时 09:29:15
English(EN) The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

新研究质疑VLM置信度校准,提出新的评估指标

一篇新发表在arXiv上的研究论文《校准置信度的幻象》(The Mirage of Calibrated Confidence)揭示,视觉语言模型(VLM)通常会报告其答案的高置信度,而这与其遵循的推理过程无关。研究发现,口头置信度在很大程度上独立于模型的内部轨迹,这意味着它不能准确反映其步骤的正确性。为解决此问题,研究人员提出了一种名为轨迹基础评分(Trajectory-Grounding Score, TGS)的新评估指标,以及一个名为TGS-Bench的基准测试套件,旨在通过比较模型在有无访问其推理路径时的置信度水平,来更好地评估VLM的可靠性。 AI

影响 强调了VLM评估中的一个关键缺陷,可能导致更可靠的模型和对其局限性的更好理解。

排序理由 发表在arXiv上的研究论文,提出了一种新的视觉语言模型评估指标。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新研究质疑VLM置信度校准,提出新的评估指标

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发表在arXiv上的研究论文,提出了一种新的视觉语言模型评估指标。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jisoo Yang, Jaeho Han, Trung X. Pham, Junyeong Kim ·

    校准置信度的海市蜃楼:视觉语言模型口头置信度的轨迹独立性

    arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence. We find that this occurs because verbalized confidence is largely traj…