A new study published on arXiv reveals significant flaws in multiple-choice Visual Question Answering (MC-VQA) benchmarks, which are commonly used to evaluate Multimodal Large Language Models (MLLMs). Researchers found that performance on these benchmarks is highly sensitive to subtle, semantically neutral changes in prompt formatting, such as option ID sets, delimiters, and separators. These formatting variations can lead to rank reversals in model performance, indicating that MC-VQA may not reliably measure multimodal reasoning capabilities. The study suggests that current MC-VQA evaluations do not adequately control for prompt format sensitivity, leading to unreliable benchmarking and motivating the development of new evaluation protocols. AI
IMPACT Highlights critical issues in current LLM evaluation, potentially impacting how multimodal model performance is assessed.
RANK_REASON Academic paper detailing research findings on AI evaluation methods. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →