PulseAugur
实时 10:58:56

LiveXiv benchmark 使用 ArXiv 论文测试多模态 AI 模型

研究人员推出了 LiveXiv,这是一个新颖的基准测试,通过从 ArXiv 论文中动态生成视觉问答对来评估大型多模态模型 (LMM)。此方法旨在防止测试数据污染,并更准确地评估模型能力。LiveXiv 自动从手稿中提取图表等内容,无需人工干预,并且高效的评估方法降低了总体成本。该基准测试已被用于测试多个开源和专有 LMM,其中手动验证的子集与自动标注相比,性能差异极小。 AI

影响 为评估多模态 AI 模型提供了一种更稳健的方法,有可能推动其在现实世界知识和推理能力方面的改进。

排序理由 该集群描述了一个用于评估 AI 模型的新基准测试,该测试在学术论文中提出。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LiveXiv benchmark 使用 ArXiv 论文测试多模态 AI 模型

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Nimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin, M. Jehanzeb Mirza, Leshem Choshen, Mikhail Yurochkin, Yuekai Sun, Assaf Arbelle, Leonid Karlinsky, Raja Giryes ·

    LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

    arXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required world knowledge to perform effectively on multiple downstream tasks. However, o…