A new research paper published on arXiv investigates the reliability of frozen hematology foundation models when faced with shifts in data acquisition pipelines. The study found that while these models achieve high accuracy on in-domain data, their performance significantly degrades when applied to data from different scanners, sites, or preparation methods. Calibration also collapses in these cross-dataset scenarios, leading to confidently incorrect predictions. The research proposes Class-Balanced Re-standardization (CBR) as a method to improve robustness and calibration, though encoder-level exceptions and residual miscalibration persist. AI
IMPACT Highlights the critical need for robust evaluation of foundation models beyond in-domain performance, especially for clinical applications.
RANK_REASON Research paper published on arXiv detailing model performance limitations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →