发表机构
Charles Sturt University(查尔斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用孟加拉国2007-2022年人口与健康调查数据,评估多种机器学习模型预测儿童发育迟缓的时间稳健性与亚组公平性,发现TabPFN和AdaBoost表现最佳,并强调时间验证与公平性评估的重要性。
AI 中文摘要
儿童发育迟缓仍然是孟加拉国的一个重大公共卫生问题,反映了受儿童、母亲、家庭、社会经济和卫生服务因素影响的长期生长失败。本研究使用了2007年至2022年具有全国代表性的孟加拉国人口与健康调查数据,开发用于人群水平儿童发育迟缓预测的机器学习模型,并评估时间稳健性和亚组公平性。纳入0-59个月且具有完整人体测量和预测变量数据的儿童。2007年、2011年和2014年调查轮次的数据用于模型开发,而2018年和2022年轮次保留作为时间测试数据集。评估了十二种特征选择方法,并使用KNN排列重要性选择的预测变量集进行最终模型评估。评估了十一种机器学习模型:十种常规算法和一种预训练的表格基础模型TabPFN。使用平衡准确率、AUROC、F1分数、Brier分数和期望校准误差评估性能。按儿童性别、居住地和社会经济状况检查亚组公平性。最终分析样本包括18,844名儿童,其中35.05%发育迟缓。在开发留出测试数据集中,TabPFN总体平衡准确率最高,为67.58%,而AdaBoost在常规模型中平衡准确率最高,为67.51%。在时间测试中,BDHS 2018中梯度提升和BDHS 2022中XGBoost的平衡准确率最高。模型性能在不同调查轮次和亚组间存在差异,强调了在公共卫生预测建模中时间验证、亚组公平性评估和透明解释的重要性。
英文摘要
Childhood stunting remains a major public health concern in Bangladesh and reflects long-term growth failure influenced by child, maternal, household, socioeconomic, and health-service factors. This study used nationally representative Bangladesh Demographic and Health Survey data from 2007 to 2022 to develop machine learning models for population-level prediction of childhood stunting and to assess temporal robustness and subgroup fairness. Children aged 0-59 months with complete anthropometric and predictor data were included. Data from the 2007, 2011, and 2014 survey rounds were used for model development, while the 2018 and 2022 rounds were retained as temporal test datasets. Twelve feature-selection approaches were assessed, and the KNN permutation importance-selected predictor set was used for final model evaluation. Eleven machine learning models were evaluated: ten conventional algorithms and one pretrained tabular foundation model, TabPFN. Performance was assessed using balanced accuracy, AUROC, F1-score, Brier score, and expected calibration error. Subgroup fairness was examined by child sex, place of residence, and socioeconomic status. The final analytic sample included 18,844 children, of whom 35.05% were stunted. In the development hold-out test dataset, TabPFN showed the highest observed balanced accuracy overall at 67.58%, while AdaBoost showed the highest observed balanced accuracy among conventional models at 67.51%. In temporal testing, the highest observed balanced accuracy was found for Gradient Boosting in BDHS 2018 and XGBoost in BDHS 2022. Model performance varied across survey rounds and subgroups, highlighting the importance of temporal validation, subgroup fairness assessment, and transparent interpretation in public health prediction modeling.
CommentsAccepted to AusDM, 15 pages