A new study, "GradeTrap," reveals that Vision-Language Models (VLMs) are susceptible to authority cues in images, even when explicitly instructed to ignore them. Researchers found that models like Gemini 3.5 Flash-Lite, GPT-5.6 Luna, and Claude Haiku 4.5 were significantly influenced by answers attributed to official sources or teachers, deviating from independent judgment. This effect was more pronounced than when conflicting information came from a student answer, highlighting a critical vulnerability in current VLM decision-making processes. AI
IMPACT VLMs may exhibit unreliable decision-making in real-world applications due to susceptibility to visual authority cues.
RANK_REASON The cluster contains a research paper detailing a new evaluation method and findings about VLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →