PulseAugur
EN
LIVE 08:52:24

New CiteVQA benchmark exposes "Attribution Hallucination" in LLMs

A new benchmark called CiteVQA has been introduced to evaluate the evidence attribution capabilities of multimodal large language models (MLLMs). Current document question-answering (Doc-VQA) evaluations only assess the final answer, overlooking instances where models might cite incorrect sources. CiteVQA requires models to provide element-level bounding-box citations alongside answers, assessing both for accuracy. The benchmark includes 1,897 questions across 711 PDFs in various domains and languages, with an automated pipeline for generating ground-truth citations. Testing revealed significant "Attribution Hallucination," where even the best-performing model achieved only 76.0% Strict Attributed Accuracy (SAA), highlighting a reliability gap in current MLLMs for high-stakes applications. AI

IMPACT Highlights a critical reliability gap in LLMs for high-stakes domains, necessitating new evaluation methods for trustworthy document intelligence.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New CiteVQA benchmark exposes "Attribution Hallucination" in LLMs

How we ranked this

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Bangrui Xu, Yuhan Wang, Bin Wang, Conghui He ·

    CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

    arXiv:2605.12882v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only approach mask…