DocVQA
PulseAugur coverage of DocVQA — every cluster mentioning DocVQA across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
SkillAA framework enhances LLM external skill integration with attribution-guided updates
A new framework called SkillAA has been developed to improve how large language models interact with external skills. This system uses a skill graph to guide the selection, repair, and validation of these skills, contra…
-
New system enhances document Q&A with visual retrieval and evidence threading
Researchers have developed a novel system for question answering on long-context documents, particularly those with visual elements like charts and infographics. The system, named VisRAG-Ret, utilizes a frozen Qwen2.5-V…
-
New methods enhance visual document question answering with adaptive retrieval and agentic restoration
Researchers have developed new methods to improve Visual Document Question Answering (DocVQA) and Knowledge-Based Visual Question Answering (KB-VQA). ViSAR introduces an adaptive retrieval technique that dynamically sel…
-
New Auditing Method Assesses Visual Token Provenance in MLLMs
A new research paper introduces a method for auditing the spatial provenance of visual tokens in multimodal large language models (MLLMs). This approach goes beyond traditional accuracy metrics to assess whether a model…
-
DeCoRAG pipeline enhances multimodal RAG for complex documents
Researchers have introduced DeCoRAG, a novel multimodal Graph RAG pipeline designed to improve complex document understanding. This new approach addresses the "Visual Attention Sink" problem, where vision-language model…
-
New framework Seer accelerates DMLLMs by up to 31x via MLP sparsity
Researchers have developed a new framework called Seer that significantly accelerates the inference speed of Diffusion Multimodal Large Language Models (DMLLMs). By analyzing the MLP activation sparsity in the first den…
-
New benchmark SynthDocBench reveals VLM failures in long-context document understanding
Researchers have introduced SynthDocBench, a novel synthetic benchmark designed to evaluate the long-context visual document understanding capabilities of vision-language models (VLMs). Unlike existing benchmarks, Synth…
-
Study finds vision-language models struggle with complex document layouts
A new study evaluates eight open-source vision-language models (VLMs) on their ability to perform Document Visual Question Answering (DocVQA) across three distinct document types: industrial documents, infographics, and…
-
Study finds visual understanding limits VLM performance on complex documents
A new study evaluates eight open-source Vision-Language Models (VLMs) on Document Visual Question Answering (DocVQA) across industrial documents, infographics, and presentation slides. The research found that while VLMs…
-
New 'Counterfeit Answers' attack targets OCR-free DocVQA models
Researchers have developed a novel adversarial attack method called "Counterfeit Answers" that can forge document content to manipulate OCR-free Document Visual Question Answering (DocVQA) models. This attack can induce…
-
SoftSkill method compresses LLM skills into compact latent controls
Researchers have developed SoftSkill, a novel method for adapting large language models to specific tasks by compressing skills into compact, continuous context objects. This approach refines a frozen backbone model wit…
-
New VQA methods enhance explainability and knowledge integration for multimodal LLMs
Researchers have developed CoExVQA, a new framework for Document Visual Question Answering (DocVQA) that enhances explainability by breaking down the reasoning process. This method first identifies relevant evidence, th…