PulseAugur
EN
LIVE 09:44:24

LiveXiv benchmark uses ArXiv papers to test multi-modal AI models

Researchers have introduced LiveXiv, a novel benchmark designed to evaluate large multi-modal models (LMMs) by dynamically generating visual question-answering pairs from ArXiv papers. This approach aims to prevent test data contamination and provide a more accurate assessment of model capabilities. LiveXiv automatically extracts content like graphs and tables from manuscripts without human intervention, and an efficient evaluation method reduces overall costs. The benchmark has been used to test several open and proprietary LMMs, with a manually verified subset showing minimal performance variance compared to automatic annotations. AI

IMPACT Provides a more robust method for evaluating multi-modal AI models, potentially driving improvements in their real-world knowledge and reasoning capabilities.

RANK_REASON The cluster describes a new benchmark for evaluating AI models, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LiveXiv benchmark uses ArXiv papers to test multi-modal AI models

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Nimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin, M. Jehanzeb Mirza, Leshem Choshen, Mikhail Yurochkin, Yuekai Sun, Assaf Arbelle, Leonid Karlinsky, Raja Giryes ·

    LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

    arXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required world knowledge to perform effectively on multiple downstream tasks. However, o…