arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于CPRD的患有多种长期疾病的老年患者住院风险预测的可扩展临床数据基础设施与机器学习对比评估

Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD

Asra Aslam, Volodymyr Chapman, Maurice M. O'Connell, Aseel S. Abuzour, Michael Abaho, Danushka Bollegala, Gary Leeming, Eduard Shantsila, Andrew Clegg, Lauren E. Walker, Iain Edward Buchan, Samuel D. Relton

arXiv 2608.29419首次发表:更新:

发表机构

School of Information, University of Sheffield; School of Medicine, Faculty of Medicine and Health, University of Leeds; Division of Informatics, University of Manchester; Academic Unit for Ageing & Stroke Research, University of Leeds; Institute of Population Health, University of Liverpool(谢菲尔德大学信息学院; 利兹大学医学与健康学院医学院; 曼彻斯特大学信息学系; 利兹大学衰老与中风研究学术单元; 利物浦大学人口健康研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究基于CPRD数据构建临床数据基础设施,对比TG-CNN、LASSO逻辑回归、随机森林预测老年多病患者住院风险,发现LASSO虽区分度非最优但校准性好,更适合临床部署。

AI 中文摘要

深度学习架构越来越多地被提出用于电子健康记录(EHR)中的患者轨迹建模,但其相比更简单、更具可解释性的模型的优势,在真实临床场景中很少受到严格的实证检验。我们提出了一套应用于CPRD Aurum中老年患者的综合患者时间线流程,纳入了通过三层自动化框架分类的260种临床疾病,该框架包含针对17种复杂疾病的专门检测逻辑。利用这一基础设施,我们以(但不限于)药物不良反应的升高风险为动机,将时间图卷积神经网络(TG-CNN)与带LASSO正则化的逻辑回归、随机森林进行基准测试,用于预测12个月全因急诊住院风险。在交叉验证中,TG-CNN的平均AUC-ROC略高于LASSO(0.712对比0.705),而在保留的测试集上,LASSO是三个模型中区分度最高的(AUC-ROC为0.733,随机森林为0.710,TG-CNN为0.702)。我们表明,仅区分度是临床部署的不完整标准:经过Platt校准后,LASSO是唯一具有可接受校准斜率的模型(0.817),而随机森林(0.759)和TG-CNN(0.391)仍存在严重的校准偏差。我们认为,最适合直接临床部署的模型是LASSO,而非区分度最高的模型。我们向机器学习和医疗保健领域的同行提供了关于数据基础设施、模型选择以及高风险决策支持中校准和可解释性价值的经验教训。

英文摘要

Deep learning architectures are increasingly proposed for patient trajectory modeling in electronic health records (EHRs), yet their advantage over simpler, more interpretable models is rarely subjected to rigorous empirical scrutiny in real-world clinical settings. We present a comprehensive patient timeline pipeline applied to elderly patients in CPRD Aurum, incorporating 260 clinical conditions classified via a three-tier automated framework including specialised detection logic for 17 complex conditions. Using this infrastructure, we benchmark Temporal Graph Convolutional Neural Networks (TG-CNN) against Logistic Regression with LASSO regularisation and Random Forests for predicting 12-month all-cause emergency hospitalisation risk, motivated by (but not filtered to) the elevated risk of adverse drug reactions. Under cross-validation, TG-CNN achieves a marginally higher mean AUC-ROC than LASSO (0.712 vs. 0.705), whereas on the held-out test set LASSO achieves the highest discrimination of three models (AUC-ROC 0.733, versus 0.710 for Random Forest and 0.702 for TG-CNN). We show, that discrimination alone is an incomplete criterion for clinical deployment: after Platt calibration, LASSO is the only model with an acceptable calibration slope (0.817), while Random Forest (0.759) and, TG-CNN (0.391) remain substantially miscalibrated. We argue that LASSO, not the highest-discriminating model, is the model best suited to direct clinical deployment. We present lessons for the machine learning and healthcare community regarding data infrastructure, model selection, and value of calibration and interpretability in high-stakes decision support.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑