Researchers have introduced MPS-Bench, a new benchmark designed to evaluate personalized safety in vision-language models (VLMs). The benchmark consists of 5,181 scenarios derived from real-world images and includes hidden user profiles to test how VLMs handle sensitive information. Evaluations of eight leading VLMs revealed that they frequently fail to defer to missing context, with none scoring above 2.6 out of 5 on personalized safety. The study identified "visual dominance" as a key issue, where visual information early in processing can override textual safety signals, leading to unsafe responses. To address this, a new method called PRISM was developed, which acts as an input monitor to predict when deferral is necessary, achieving a high AUC of 0.978. AI
IMPACT Highlights critical safety limitations in VLMs, potentially influencing future development and deployment strategies for multimodal AI.
RANK_REASON The cluster contains an academic paper detailing a new benchmark and method for evaluating AI safety. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →