Two new research papers highlight critical issues with using public datasets for training medical AI models, particularly for chest radiograph analysis. The first paper, focusing on vision-language models, found that agreement with institutional reference standards varied significantly across sites and findings, suggesting a need for site-specific re-evaluation. The second paper introduced a framework to audit repository labels against expert annotations, revealing near-zero agreement for cardiomegaly in the MIMIC-CXR dataset and emphasizing the unreliability of repository-derived labels as ground truth. AI
IMPACT Highlights significant limitations in current medical AI training data, potentially impacting the reliability and deployment of diagnostic tools.
RANK_REASON Two academic papers published on arXiv presenting new research findings and frameworks.
- arXiv
- Beta-Binomial empirical-Bayes estimator
- CORE Recommender
- Hugging Face
- logistic regression model
- vision-language model
- cardiomegaly
- DenseNet121
- MIMIC-CXR
- Repository Supervision Auditing
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →