Researchers have introduced VeriOCRBench, a new benchmark designed to evaluate the task verification capabilities of Multimodal Large Language Models (MLLMs) in optical character recognition (OCR) scenarios. This benchmark addresses the issue of "blind compliance" where models often assume tasks are valid, even when faced with illegible text, contradictory premises, or missing information. VeriOCRBench includes 1,800 samples with injected invalid tasks across various domains and verification dimensions, aiming to measure a model's ability to determine if a task is executable before attempting to answer. Evaluations of 15 leading MLLMs revealed significant reliability gaps, including persistent blind compliance and issues with over-refusal. AI
IMPACT Highlights a critical reliability gap in current OCR reasoning systems, potentially driving development of more robust and trustworthy MLLMs for document understanding.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →