发表机构
Stony Brook University; Stony Brook Medicine(石溪大学; 石溪医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用All of Us数据,证明患者报告调查数据可显著提升阿片类药物使用障碍(OUD)的预测性能,其中24个月LightGBM模型PR-AUC提升至0.6603,并强调调查可用性的重要性。
AI 中文摘要
电子健康记录(EHRs)可能无法完整捕捉与阿片类药物使用障碍(OUD)相关的患者报告因素。我们评估了调查数据是否能改善对267,747名有记录的阿片类药物暴露的All of Us参与者中首次记录的OUD诊断的预测,其中包括15,287例OUD病例。我们使用逻辑回归、随机森林、XGBoost、LightGBM、多层感知器、LSTM、GRU和Transformer,在6个月、12个月和24个月的回看窗口内比较了仅使用EHR和EHR+调查模型。调查数据增强在所有24个模型-窗口组合中将PR-AUC提高了0.0087-0.0505;最佳的24个月LightGBM模型从0.6219提高到0.6603。调查覆盖率随窗口延长而增加,并因OUD状态而异(24个月:OUD阳性为21.7%,OUD阴性为60.7%)。排列分析将调查特征列为24个月时两种评估模型中第二重要的信息域。患者报告的数据提供了超越结构化EHR的互补预测信号,同时强调了调查可用性的重要性。
英文摘要
Electronic health records (EHRs) may incompletely capture patient-reported factors associated with opioid use disorder (OUD). We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases. We compared EHR-only and EHR+survey models across 6-, 12-, and 24-month look-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer. Survey augmentation improved PR-AUC across all 24 model-window combinations by 0.0087-0.0505; the best 24-month LightGBM model improved from 0.6219 to 0.6603. Survey coverage increased with longer windows and differed by OUD status (24 months: 21.7% OUD-positive vs. 60.7% OUD-negative). Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models. Patient-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability.
Comments10 pages; submitted to the AMIA 2027 Amplify Informatics Summit