AI 中文总结
研究针对电子健康记录自回归基础模型预训练方法存在的问题,提出患者采样方法,通过控制训练信号分布来构建序列,实验表明该方法在真实EHR数据上提升了性能,凸显训练和验证序列构建对相关模型的重要性。
AI 中文摘要
用于电子健康记录(EHR)的自回归基础模型通常继承语言建模的预训练方法,将患者轨迹连接成单个令牌流并从中采样窗口。在EHR数据中,这种选择很重要:窗口可能混合多个患者,记录较长的患者贡献更多优化更新,可能引入偏差。我们提出患者采样,一种预训练序列构建方法,可控制训练信号在患者间的分布。我们将此方法与标准方法(全局流)比较。结果表明,具有可控加权的随机患者采样提高了对真实EHR数据的性能。在MIMIC-IV v2.2和v3.1的下游临床任务中,患者采样比全局流基线提高了宏AUROC和AUPRC。这些结果表明训练和验证序列构建是自回归EHR基础模型重要且未充分探索的设计选择。
英文摘要
Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into a single token stream and windows are sampled from that stream. In EHR data, this choice is consequential: windows may mix multiple patients, and patients with longer records contribute more optimization updates, potentially introducing bias. We propose Patient Sampling, a pretraining sequence-construction method that allows us to control how training signal is distributed across patients. We compare this method to the standard approach, which we refer to as Global Stream. We show that stochastic Patient Sampling with controllable weighting improves performance on real-world EHR data. Across downstream clinical tasks on MIMIC-IV v2.2 and v3.1, Patient Sampling improves Macro AUROC and AUPRC over the Global Stream baseline. These results identify training and validation sequence construction as important and underexplored design choices for autoregressive EHR foundation models.
CommentsAccepted at the Workshop on Structured Data for Health at the 43rd International Conference on Machine Learning (ICML 2026). 7 pages, 4 figures
Journal refICML 2026 Workshop on Structured Data for Health (SD4H)