AI 中文总结
研究利用结构化国家层面每周数据,通过可解释机器学习框架预测呼吸疾病率和空气质量状况,比较多种回归和分类模型,经SHAP值解释及亚组分析发现PM2.5是主要预测因子,表明仅模型准确性不足以进行气候 - 健康预测,可解释模型有重要作用。
AI 中文摘要
空气污染和气候相关压力源对呼吸健康愈发重要,尤其在环境暴露和医疗能力不平等的地区。本研究评估一个可解释机器学习框架,使用结构化国家层面每周数据预测呼吸疾病率和空气质量状况。考虑两项监督学习任务,通过嵌套交叉验证比较九个回归模型和九个分类模型。用SHAP值进行模型解释,并按收入水平和地理区域进行亚组分析。结果表明,PM2.5浓度是呼吸疾病率的主要预测因子,线性和正则化线性模型回归性能最强。空气质量分类中,包含PM2.5时模型平衡准确率高,去除后性能大幅下降。SHAP分析显示,无PM2.5时,人均GDP、降水和医疗可及性等社会经济和气象变量影响更大。亚组分析表明,各收入组总体回归误差相似,但PM2.5在中低收入国家对预测贡献更大。这些结果表明,仅模型准确性不足以进行气候 - 健康预测。可解释模型有助于识别主要污染相关信号,测试结果是否依赖关键污染物变量,以及显示预测模式在社会经济群体间是否不同。
英文摘要
Air pollution and climate-related stressors are increasingly important concerns for respiratory health, especially in settings with unequal environmental exposure and healthcare capacity. This study evaluates an interpretable machine learning framework for predicting respiratory disease rates and air-quality status using structured country-level weekly data. Two supervised learning tasks were considered: regression of respiratory disease rate per 100,000 population and binary classification of air-quality status. Nine regression models and nine classification models were compared using nested cross-validation. Model interpretation was conducted using SHAP values, and subgroup analysis was performed across income levels and geographic regions. The results showed that PM2.5 concentration was the dominant predictor of respiratory disease rate, with linear and regularized linear models achieving the strongest regression performance. For air-quality classification, models achieved high balanced accuracy when PM2.5 was included, but performance decreased substantially when PM2.5 was removed, indicating strong dependence on pollutant-related information. SHAP analysis showed that, without PM2.5, socioeconomic and meteorological variables such as GDP per capita, precipitation, and healthcare access became more influential. Subgroup analysis showed similar aggregate regression error across income groups, but PM2.5 contributed more strongly to predictions in lower-middle-income countries. These results show that model accuracy alone is not sufficient for climate-health prediction. Interpretable models can help identify dominant pollution-related signals, test whether results depend on key pollutant variables, and show whether prediction patterns differ across socioeconomic groups.