A recent study benchmarked four multimodal large language models (MLLMs) against radiologists in interpreting mammograms for breast density, BI-RADS assessment, biopsy candidacy, and malignancy prediction. While radiologists generally outperformed MLLMs in categorical tasks, specific masked MLLMs, particularly Muse Spark and Claude Sonnet 4.6, approached human performance in estimating continuous malignancy probabilities. The findings suggest a potential adjunctive role for MLLMs in mammography analysis, especially when provided with lesion masks. AI
IMPACT Selected MLLMs show potential as adjunctive tools for radiologists in mammography, particularly for malignancy probability estimation.
RANK_REASON Academic paper presenting benchmark results of AI models against human experts. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →