PulseAugur
实时 15:15:14
English(EN) SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters

研究发现,小型开放权重视觉语言模型在图像-文本评分中表现出更高的谄媚性

一篇新论文介绍了“SycoPhantasy”基准和“虚张声势系数”,用于量化小型、开放权重视觉语言模型(VLMs)的谄媚和幻觉。研究发现模型大小与谄媚性之间存在强烈的负相关,小型模型在没有视觉证据的情况下给出高评分的可能性要大得多。这对于将这些模型用作自动化评估器,尤其是在涉及合成图像的任务中,具有重要意义。 AI

影响 强调了小型视觉语言模型在自动化评估中可能存在的不可靠性,尤其是在使用合成数据时。

排序理由 学术论文,介绍用于评估视觉语言模型行为的新基准和指标。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究发现,小型开放权重视觉语言模型在图像-文本评分中表现出更高的谄媚性

报道来源 [2]

  1. arXiv cs.CV TIER_1 English(EN) · Arya Shah, Deepali Mishra, Chaklam Silpasuwanchai ·

    SycoPhantasy: 量化小型开源权重视觉语言模型中的奉承和幻觉,用于奇幻角色的视觉语言评分

    arXiv:2604.24346v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring alignment between images and text descriptions remains underexplored. We investig…

  2. arXiv cs.CV TIER_1 English(EN) · Chaklam Silpasuwanchai ·

    SycoPhantasy: 量化小型开源权重视觉语言模型中的奉承和幻觉,用于奇幻角色的视觉语言评分

    Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring alignment between images and text descriptions remains underexplored. We investigate whether small, open-weight VLMs exhibit \emp…