PulseAugur
实时 10:11:38
English(EN) CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

新的CAPEval基准将图像字幕质量解耦为覆盖度和精确度

研究人员推出了一种新的基准CAPEval,旨在通过将图像字幕解耦为两个不同的属性:覆盖度和精确度来评估图像字幕。覆盖度衡量字幕对视觉内容的描述程度,而精确度则评估字幕中声明的事实准确性。实验表明,覆盖度是多模态理解任务性能的更强预测因子,而精确度对文本到图像生成任务的影响更大。这种方法提供了对字幕质量更细致的评估,并为针对特定下游应用优化字幕生成器提供了指导。 AI

影响 提供了对图像字幕质量更细粒度的评估,有助于针对特定的多模态任务优化字幕模型。

排序理由 该条目描述了一个用于评估图像字幕的新学术基准。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CAPEval基准将图像字幕质量解耦为覆盖度和精确度

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang ·

    CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    arXiv:2608.02589v1 Announce Type: new Abstract: Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct proper…