Researchers have introduced BEAR-Bench, a new benchmark designed to evaluate the reasoning capabilities of multimodal large language models (MLLMs) on complex, text-dense documents in both English and Russian. The benchmark addresses limitations in existing evaluations, which often focus on simple information extraction or are heavily biased towards English and Chinese. BEAR-Bench includes 1000 human-annotated questions derived from business and scientific texts, and initial evaluations of 16 MLLMs, including Gemini-3.1 Pro and Qwen3.5-397B, reveal significant room for improvement even in top-performing models. The study also compares existing hallucination detection methods using the benchmark's outputs. AI
IMPACT This benchmark could drive improvements in multimodal LLM reasoning for professional documents across multiple languages.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →