arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

2型糖尿病风险预测的机器学习综合评估:大规模外部验证与公平性分析

Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis

Rajveer Singh Pall, Sameer Yadav, Siddharth Bhalerao, Sourabh Sahu, Ritu Ahluwalia, Bhaskar Awadhiya

arXiv 2607.16253首次发表:更新:

AI 中文总结

研究针对2型糖尿病风险预测模型外部测试和公平性评估不足问题,开发多维框架,用NHANES数据训练XGBoost模型并在BRFSS上验证,发现存在性能损失和公平性偏差,确定主要风险驱动因素,强调需公平性感知、年龄分层的部署策略。

AI 中文摘要

基于机器学习的2型糖尿病风险预测模型内部验证结果良好,但因外部测试和公平性评估不足,在实际应用中效果不佳。我们开发了一个多维框架,在全国代表性人群上评估歧视、校准、可解释性和算法公平性。使用八个非实验室预测变量在2015 - 2020年的美国国家健康与营养检查调查(NHANES,n = 15685)上训练XGBoost模型。在2020 - 2022年的美国国家健康访问调查(BRFSS,n = 1285783)上进行外部验证。内部验证显示有良好的区分能力(AUC = 0.794,95% CI 0.788 - 0.800),外部验证性能下降(AUC = 0.717,相对下降:-9.7%,p < 0.001)。公平性分析揭示了严重偏差,校准显示风险高估。SHAP分析确定了年龄、BMI和身体活动是主要风险驱动因素。高糖尿病风险人群算法性能最差,强调临床使用前需要公平性感知、年龄分层的部署策略。

英文摘要

Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment. We developed a multi-dimensional framework evaluating discrimination, calibration, interpretability, and algorithmic fairness on nationally representative populations. An XGBoost model was trained on NHANES 2015-2020 (n=15,685) using eight non-laboratory predictors: age, sex, race/ethnicity, BMI, smoking status, physical activity, history of heart attack, and history of stroke. External validation was performed on BRFSS 2020-2022 (n=1,285,783) under realistic distribution shift. Internal validation showed good discrimination (AUC=0.794, 95% CI 0.788-0.800), with performance loss on external validation (AUC=0.717, relative decrease: -9.7%, p<0.001). Fairness analysis revealed severe bias: elderly adults (>=60) showed AUC=0.607 vs 0.742 for young adults (difference=0.135, p<0.001); obese individuals showed AUC=0.698 vs 0.735 for normal weight (difference=0.037, p<0.001). Gender showed comparable performance (male=0.723 vs female=0.712, p=0.142). Calibration revealed risk overestimation (Brier score=0.123). SHAP analysis identified age, BMI, and physical activity as primary risk drivers. Populations with highest diabetes risk receive the worst algorithmic performance, underscoring the need for fairness-aware, age-stratified deployment strategies before clinical use.

CommentsAccepted and published at the IEEE EDS Technically Sponsored International Conference on Intelligent Processing, Hardware, Electronics, and Radio Systems (CIPHER-2026), 13-15 Feb 2026, NIT Jalandhar, India (IEEE Conference Record #70417, Paper ID: 155). 8 pages, 4 figures, 3 tables

Journal refProc. 2026 Int. Conf. on Intelligent Processing, Hardware, Electronics and Radio Systems (CIPHER), Jalandhar, India, Feb. 2026

DOI:10.1109/CIPHER70417.2026.11523789

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑