Researchers have introduced DocHop, a new benchmark designed to evaluate the multi-hop reasoning capabilities of Multimodal Large Language Models (MLLMs) when dealing with information-dense documents. Unlike existing benchmarks that assess chart and text understanding in isolation, DocHop requires models to use textual context to select, interpret, and aggregate data from charts. The benchmark, comprising 2,074 examples across six categories, was generated using a stochastic pipeline to control reasoning depth and visual density. Experiments reveal a significant performance gap between human annotators (over 90% accuracy) and the best-performing models (62.83%), with performance degrading as reasoning complexity increases. AI
IMPACT DocHop aims to push MLLM capabilities in complex document understanding, potentially leading to more sophisticated AI assistants for research and data analysis.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →