Researchers have introduced LiveXiv, a novel benchmark designed to evaluate large multi-modal models (LMMs) by dynamically generating visual question-answering pairs from ArXiv papers. This approach aims to prevent test data contamination and provide a more accurate assessment of model capabilities. LiveXiv automatically extracts content like graphs and tables from manuscripts without human intervention, and an efficient evaluation method reduces overall costs. The benchmark has been used to test several open and proprietary LMMs, with a manually verified subset showing minimal performance variance compared to automatic annotations. AI
IMPACT Provides a more robust method for evaluating multi-modal AI models, potentially driving improvements in their real-world knowledge and reasoning capabilities.
RANK_REASON The cluster describes a new benchmark for evaluating AI models, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Leshem Choshen
- LiveXiv
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →