PulseAugur
EN
LIVE 09:42:49

New VTS method reveals MLLMs struggle with visual instructions

Researchers have introduced Visualized Task Semantics (VTS), a novel method to evaluate how well multimodal large language models (MLLMs) understand instructions presented visually within images. When questions were moved into the image across six MLLMs and four benchmarks, accuracy dropped by an average of 17.8 points, indicating a significant gap in how models process visual instructions compared to text. To address this, the team developed prompt-region grounding, a technique that aligns question regions with typed semantics, improving VTS accuracy from 58.0% to 66.3% without requiring OCR or region metadata at inference. AI

IMPACT Highlights a critical limitation in MLLMs' ability to follow visual instructions, potentially guiding future model development towards better multimodal reasoning.

RANK_REASON Research paper introducing a new method and benchmark for evaluating MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New VTS method reveals MLLMs struggle with visual instructions

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang ·

    When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

    arXiv:2608.04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text, leaving it unclear whether models use the same instr…