A new benchmark called CiteVQA has been introduced to evaluate the evidence attribution capabilities of multimodal large language models (MLLMs). Current document question-answering (Doc-VQA) evaluations only assess the final answer, overlooking instances where models might cite incorrect sources. CiteVQA requires models to provide element-level bounding-box citations alongside answers, assessing both for accuracy. The benchmark includes 1,897 questions across 711 PDFs in various domains and languages, with an automated pipeline for generating ground-truth citations. Testing revealed significant "Attribution Hallucination," where even the best-performing model achieved only 76.0% Strict Attributed Accuracy (SAA), highlighting a reliability gap in current MLLMs for high-stakes applications. AI
IMPACT Highlights a critical reliability gap in LLMs for high-stakes domains, necessitating new evaluation methods for trustworthy document intelligence.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- Attribution Hallucination
- CiteVQA
- Doc-VQA
- Dongsheng Manchu and Mongol Ethnic Township
- finance
- Gemini 3.1-pro-preview
- law
- medicine
- Multimodal Large Language Models
- Strict Attributed Accuracy
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →