A new pretraining method called Patient Sampling has been developed for autoregressive foundation models used with electronic health records (EHRs). This method addresses biases that can arise from standard language modeling approaches, where patient data is concatenated and windows may mix multiple patients. By controlling how training signals are distributed, Patient Sampling improves performance on downstream clinical tasks, showing enhanced Macro AUROC and AUPRC scores on MIMIC-IV datasets compared to the Global Stream baseline. The research highlights sequence construction as a critical, yet often overlooked, design choice for EHR foundation models. AI
IMPACT This new method could lead to more accurate and less biased AI models for analyzing electronic health records, improving clinical decision-making.
RANK_REASON Academic paper detailing a new method for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →