MMLongBench-Doc
PulseAugur coverage of MMLongBench-Doc — every cluster mentioning MMLongBench-Doc across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New system enhances document Q&A with visual retrieval and evidence threading
Researchers have developed a novel system for question answering on long-context documents, particularly those with visual elements like charts and infographics. The system, named VisRAG-Ret, utilizes a frozen Qwen2.5-V…
-
New MCite-RL framework enhances multimodal RAG with citation-enhanced reinforcement learning
Researchers have developed MCite-RL, a new framework designed to improve the reliability of multimodal Retrieval-Augmented Generation (RAG) systems. This approach uses an agentic reinforcement learning method to enhance…
-
D2-ScaleAgent framework enhances long document understanding
Researchers have introduced D2-ScaleAgent, a novel framework designed to enhance the understanding of long and visually rich documents. This agentic system employs a dual-dimensional scaling paradigm, dynamically adjust…
-
New Trident method enhances multimodal QA for long documents
Researchers have developed a new method called Trident to improve multimodal question answering over long documents. Trident consists of two components: Trident-R, an LLM reranker that creates structured semantic record…
-
New frameworks tackle long and evolving document understanding challenges
Researchers have developed new frameworks to tackle the challenges of understanding long and evolving documents. InSight-doc, an agentic visual perception framework, adaptively allocates visual resolution to improve acc…
-
New benchmarks and frameworks tackle extra-long document understanding
Researchers have introduced two new frameworks for improving the ability of large language models to understand and answer questions from very long documents. DocTrace focuses on creating a traceable evidence graph to s…
-
New VLD-RAG framework enhances AI's ability to process long, visual documents
Researchers have developed VLD-RAG, a novel agentic framework designed for retrieval-augmented generation over long, visually-rich documents. This system constructs a multimodal index that preserves page layout and inco…
-
New TAP-RAG framework improves multimodal QA on long documents
Researchers have introduced TAP-RAG, a novel framework designed to enhance multimodal question answering over long documents. This system employs a Task-Aware Policy Controller (TAPC) that analyzes queries to determine …
-
New benchmark SynthDocBench reveals VLM failures in long-context document understanding
Researchers have introduced SynthDocBench, a novel synthetic benchmark designed to evaluate the long-context visual document understanding capabilities of vision-language models (VLMs). Unlike existing benchmarks, Synth…
-
New framework learns to dynamically orchestrate AI retrievers for document reasoning
Researchers have developed a novel framework for multimodal document reasoning agents that learns to dynamically orchestrate various retrieval methods. This failure-driven evolution approach allows a meta-agent to adapt…
-
MAGE-RAG framework enhances multimodal QA for long documents
Researchers have introduced MAGE-RAG, a novel framework designed to improve multimodal question answering for long documents. This system constructs an adaptive graph of evidence, incorporating text, images, tables, and…
-
EviProp method improves long document retrieval with graph diffusion
Researchers have developed EviProp, a novel method for retrieving relevant pages from long, visually rich documents. Unlike existing approaches that score pages independently, EviProp models documents as multimodal Chun…
-
New CDS method advances multimodal document question answering
Researchers have developed a new retrieval method called Constrained Dominant Sets (CDS) for multimodal document question answering. This technique addresses limitations in current systems that struggle with long docume…
-
MARDoc framework enhances multimodal long document QA with structured memory
Researchers have introduced MARDoc, a novel framework designed to improve question answering for long, multimodal documents. This system utilizes three specialized agents: an Explorer for retrieval, a Refiner for proces…
-
LLMs with Vision Capabilities Tested Against OCR for Document QA
A benchmark compared vision-capable large language models against OCR-based pipelines for question-answering on long, image-heavy documents. The evaluation used 30 PDFs from the MMLongBench-Doc dataset, assessing the mo…