Researchers have introduced Visualized Task Semantics (VTS), a novel method to evaluate how well multimodal large language models (MLLMs) understand instructions presented visually within images. When questions were moved into the image across six MLLMs and four benchmarks, accuracy dropped by an average of 17.8 points, indicating a significant gap in how models process visual instructions compared to text. To address this, the team developed prompt-region grounding, a technique that aligns question regions with typed semantics, improving VTS accuracy from 58.0% to 66.3% without requiring OCR or region metadata at inference. AI
IMPACT Highlights a critical limitation in MLLMs' ability to follow visual instructions, potentially guiding future model development towards better multimodal reasoning.
RANK_REASON Research paper introducing a new method and benchmark for evaluating MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →