基于人群健康的机器学习揭示社会心理因素与慢性肾脏病之间的关联
Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease
浏览论文内容
中文总结 AI 辅助
本研究结合大规模人群健康调查数据与定制堆叠集成机器学习模型,识别出定期体检、年龄等CKD关键预测因子,为CKD风险分层提供了可解释框架。
中文摘要 AI 辅助
慢性肾脏病(CKD)进展隐匿,严重损害生活质量,因此早期检测对改善患者结局至关重要。我们开展了一项两部分研究,结合大规模远程医疗数据与先进机器学习技术,以同时对自我报告的CKD状态进行分类并识别疾病的关键驱动因素。我们使用来自行为风险因素监测系统(BRFSS 2021:438693个样本;BRFSS 2019:418268个样本)和国家健康访谈调查(NHIS 2021:29482个样本;NHIS 2020:31568个样本)的选定特征,采用9种最先进的插补方法处理缺失数据,并通过采样策略缓解类别不平衡问题。我们定制的堆叠集成模型取得了72.56-76.12%的平衡准确率,对应的AUROC分数为79.59-82.29%。SHapley加性解释(SHAP)分析经临床审查后,突出了关键预测因子,包括定期体检、年龄、血压和心理健康压力指标。这些发现为CKD风险分层提供了稳健且可解释的框架,并为其相关因素提供了可操作的见解。
英文摘要
Chronic kidney disease (CKD) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes. We present a two-part study that combines large-scale telehealth data with advanced machine learning to both classify self-reported CKD status and identify key drivers of disease. Using selected features from the Behavioral Risk Factor Surveillance System (BRFSS 2021: 438,693 samples; BRFSS 2019: 418,268 samples) and the National Health Interview Survey (NHIS 2021: 29,482 samples; NHIS 2020: 31,568 samples), we addressed missing data with nine state-of-the-art imputation methods and mitigated class imbalance via sampling strategies. Our customized stacked ensemble model achieved balanced accuracy of 72.56-76.12%, with corresponding AUROC scores of 79.59-82.29%. SHapley Additive exPlanations (SHAP) analysis, followed by clinical review, highlighted critical predictors, including regular medical check-ups, age, blood pressure, and indicators of mental health stress. These findings deliver a robust and interpretable framework for CKD risk stratification and provide actionable insights into its associated factors.