A new study published on arXiv investigates the robustness of multimodal large language models (MLLMs) when presented with conflicting information across text and image modalities. The research found that MLLMs are not consistently robust, with models favoring image-based evidence over text when contradicting parametric knowledge. Furthermore, when both text and image evidence are provided together, the model's preference appears arbitrary, influenced by input order, model, and dataset. This instability can degrade performance in multimodal RAG systems and be exploited by adversarial attacks. While simple techniques like prompting and steering were ineffective, supervised fine-tuning (SFT) showed moderate success in mitigating this brittleness, highlighting a need for greater attention to this fundamental inconsistency during model training. AI
IMPACT Highlights a critical vulnerability in multimodal LLMs that could impact RAG systems and adversarial robustness, necessitating further research into training methods.
RANK_REASON Research paper published on arXiv detailing findings about MLLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →