发表机构
University of Cambridge; Stanford University; University of Exeter(剑桥大学; 斯坦福大学; 埃克塞特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出EHRFlow,一种基于多边际流匹配的框架,通过条件化患者病史来建模连续轨迹,在合成和真实数据上提升临床代码预测与潜在状态恢复,并能在反事实模拟中近似干预效果。
AI 中文摘要
电子健康记录提供了对潜在患者状态的稀疏观测,这些状态随时间连续演化。近年来的自回归模型以临床病史为条件,将未来事件预测为离散观测序列。相反,多边际流匹配提供了一种连续时间公式,但使用多个观测来监督训练路径本身并不能让学习到的动态访问先前的患者病史。我们引入了EHRFlow,一个多边际流匹配框架,它以编码的患者病史为条件,从而使未来动态依赖于患者先前的临床轨迹。我们提出的框架适应不规则的观测时间,并支持在任意时间范围内进行预测。在受控的合成基准测试中,EHRFlow改善了临床代码预测和潜在状态恢复。在包含超过一百万患者的真实世界临床数据集上,包括一个独立的外部验证队列,EHRFlow在时间范围平均的top-5临床代码准确性上优于自回归和与历史无关的流匹配基线。最后,在受控的反事实模拟中,条件引导近似了抗高血压干预的已知效果,而无需训练特定任务的结果模型。
英文摘要
Electronic health records provide irregular observations of latent patient states that evolve continuously over time. Recent autoregressive models condition on clinical histories to forecast future events as sequences of discrete observations. Conversely, multi-marginal flow matching provides a continuous-time formulation, but using multiple observations to supervise training paths does not itself give the learned dynamics access to preceding patient history. We introduce EHRFlow, a multi-marginal flow-matching framework that conditions on encoded patient history, thereby allowing future dynamics to depend on the patient's prior clinical trajectory. Our proposed framework accommodates irregular observation times and supports forecasting at arbitrary horizons. Across controlled synthetic benchmarks, EHRFlow improves clinical-code forecasting and latent-state recovery. On real-world clinical datasets comprising more than one million patients, including an independent external validation cohort, EHRFlow improves horizon-averaged top-5 clinical-code accuracy over autoregressive and history-independent flow-matching baselines. Finally, in a controlled counterfactual simulation, conditional guidance approximates the known effect of an antihypertensive intervention without training a task-specific outcome model.