发表机构
Baylor College of Medicine(贝勒医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用儿科电子健康记录训练TEDDY模型,对纵向诊断轨迹和就诊时间建模,在疾病发病预测任务中表现出色,能支持广泛、罕见疾病和长期风险预测,无需大规模人口数据或数十亿参数模型。
AI 中文摘要
儿科电子健康记录记录了具有发育结构的临床轨迹,但其在生成式医疗基础模型中的潜力尚未得到充分探索。本文提出了TEDDY(青少年疾病时间事件解码器),这是一个拥有184万个参数的解码器变压器,在一家儿科机构的约160万名儿童的约7300万个ICD-10诊断数据上进行训练。TEDDY对纵向诊断轨迹和就诊时间进行建模。预测在就诊代码显示之前进行,限于首次出现,并与性别和年龄匹配的对照组进行评估。在跨越16个ICD-10章节的797个疾病发病预测任务中,TEDDY的中位数AUC为72.0%,在96%-99%的任务上优于相同数据的DenseNet(50.0%)、CNN(57.2%)、RNN(60.1%)和LSTM(62.7%)基线。性能在性别和年龄上保持一致,在低患病率诊断中最强;225种最罕见疾病中的202种(90%)的95%置信区间高于随机水平。在首次记录诊断前两年多仍可检测到预测信号,在无限制分析中的中位数AUC为59.7%,在固定队列敏感性分析中为64.4%。在哮喘和注意力缺陷多动障碍基准测试中,AUC分别为79.3%和84.7%,而最强的比较器(包括一个大三数量级的通用语言模型)分别为62.7%和71.7%。就诊时间预测在365天内的平均绝对受限平均生存时间误差为3.0天,尽管中位数和长尾返回间隔仍存在校准错误。这些结果共同表明,儿科诊断历史可作为紧凑生成模型的基础,支持广泛、罕见疾病和长期风险预测,而无需大规模人口数据或数十亿参数模型。
英文摘要
Pediatric electronic health records capture developmentally structured clinical trajectories, yet their potential for generative healthcare foundation models remains largely unexplored. Here we present TEDDY (Temporal Event Decoder for Disease in Youth), a 1.84-million-parameter decoder transformer trained on approximately 73 million ICD-10 diagnoses from 1.6 million children at a single pediatric institution. TEDDY models longitudinal diagnosis trajectories and visit timing. Predictions were made before visit codes were revealed, limited to first occurrences, and evaluated against sex- and age-matched controls. Across 797 disease-onset prediction tasks spanning 16 ICD-10 chapters, TEDDY achieved a median AUC of 72.0%, outperforming same-data DenseNet (50.0%), CNN (57.2%), RNN (60.1%), and LSTM (62.7%) baselines on 96-99% of tasks. Performance held across sex and age and was strongest among lower-prevalence diagnoses; 202 of the 225 rarest conditions (90%) had 95% confidence intervals above chance. Predictive signal remained detectable more than two years before first recorded diagnosis, with median AUCs of 59.7% in the unrestricted analysis and 64.4% in a fixed-cohort sensitivity analysis. In asthma and attention-deficit/hyperactivity disorder benchmarks, AUCs were 79.3% and 84.7%, compared with 62.7% and 71.7% for the strongest comparators, including a general-purpose language model three orders of magnitude larger. Visit-timing predictions had a 3.0-day mean absolute restricted mean survival-time error over 365 days, although median and long-tail return intervals remained miscalibrated. Together, these results establish pediatric diagnostic histories as a substrate for compact generative models supporting broad, rare-disease, and long-horizon risk forecasting without population-scale data or billion-parameter models.