Researchers have developed ModaLens, a new audit method to measure the sensitivity of vision-language models (VLMs) in medical contexts. Using the MedGemma-27B model on MIMIC-CXR data, the study found that when radiology reports are available, the model's reliance on image data decreases significantly. Specifically, the model's answers changed only 4.26% of the time when reports were present, compared to 20.94% when reports were absent, indicating that the textual report heavily influences the model's output, potentially at the expense of visual interpretation. AI
IMPACT Highlights potential over-reliance on text in medical VLMs, suggesting a need for improved image-grounding in clinical applications.
RANK_REASON The cluster contains an academic paper detailing a new methodology and experimental results for evaluating AI models.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →