Researchers have introduced CrossModalQA, a new benchmark designed to evaluate multimodal large language models (MLLMs) in retrieval-augmented generation (RAG) tasks. This benchmark addresses limitations in existing systems by focusing on open-domain evidence discovery and complex multi-hop reasoning across both text and images. CrossModalQA comprises over 1,800 question-answer pairs derived from Wikipedia and Wikimedia Commons, requiring models to perform multi-hop, cross-modal reasoning, with an average depth of 3.50 hops. Initial experiments indicate that current MLLMs struggle with complete evidence retrieval, and performance is significantly impacted by incomplete or distracting context. AI
IMPACT This benchmark will push the development of more robust multimodal AI systems capable of complex reasoning and evidence retrieval.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →