面向宽带接入差距的可解释机器学习:普查 tract 级预测与基于 SHAP 的因素分析
Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor Profiling
浏览论文内容
中文总结 AI 辅助
该研究提出可解释机器学习框架,以普查 tract 为单位预测美国宽带接入差距,通过LightGBM与SHAP分析识别关键影响因素,筛选精度优于仅用收入的方法,时间稳定性良好。
中文摘要 AI 辅助
美国已通过《基础设施投资和就业法案》拨款约650亿美元用于宽带扩展,但针对这些投资的循证方法仍有待开发。本文提出一种可解释机器学习框架,用于分析美国全国83359个普查 tract 层面的宽带接入差距。利用2022年美国社区调查(ACS)的65个社会经济、人口统计和基础设施特征,我们在空间五折交叉验证下训练LightGBM模型,获得R²=0.533、斯皮尔曼相关系数rho=0.763;州保留交叉验证(51折)证实模型具有泛化性(R²=0.525)。TreeSHAP分析确定收入和教育是主导因素组(工程化交互项吸收了其组成特征的归因),基于SHAP的聚类揭示了三种探索性因素概况:连接良好的中等水平(约4.9万个tract)、可负担性受限的严重水平(约2.1万个tract)、农村老年群体(约1.3万个tract)。作为筛选工具,基于机器学习的tract选择在前10%的tract中捕获了38.0%的总接入差距,而仅使用收入的启发式方法为35.2%(提升2.8个百分点,p<0.002,县-块自举检验);在遗憾减少方面,该模型缩小了仅使用收入与最优选择之间剩余差距的19%。主要贡献是tract层面的因素分解:SHAP确定哪些特征组(收入/教育、农村性、年龄)与每个tract的预测差距最密切相关,并为差异化调查提供依据。时间稳定性检验(用2017年ACS训练,预测2022年ACS且无调查年份重叠)证实排名稳定性(rho=0.784,注意超参数在2022年数据上调优)。
英文摘要
The United States has allocated approximately $65 billion through the Infrastructure Investment and Jobs Act for broadband expansion, yet evidence-based methods for targeting these investments remain underdeveloped. This paper presents an explainable machine learning framework for profiling broadband adoption disparities at census-tract granularity across 83,359 tracts nationwide. Using 65 socioeconomic, demographic, and infrastructure features derived from the American Community Survey 2022, we train a LightGBM model under spatial five-fold cross-validation, achieving R^2 = 0.533 and Spearman rho = 0.763; state-held-out cross-validation (51 folds) confirms generalization (R^2 = 0.525). TreeSHAP analysis identifies income and education as the dominant factor group (with the engineered interaction term absorbing attribution from its constituent features), and SHAP-based clustering reveals three exploratory factor profiles: Well-Connected Moderate (~49K tracts), Affordability-Limited Severe (~21K tracts), and Rural-Elderly (~13K tracts). As a screening tool, ML-based tract selection captures 38.0% of the total adoption gap within the top 10% of tracts versus 35.2% for income-only heuristics (+2.8 pp, p < 0.002, county-block bootstrap); in regret-reduction terms, the model closes 19% of the remaining gap between income-only and oracle selection. The primary contribution is the per-tract factor decomposition: SHAP identifies which feature groups (income/education, rurality, age) are most strongly associated with each tract's predicted gap, and informs differentiated investigation. A temporal stability check, training on ACS 2017 and predicting ACS 2022 with zero survey-year overlap, confirms ranking stability (rho = 0.784, noting hyperparameters tuned on 2022 data).