发表机构
Bowdoin College; Pontifical Catholic University of Campinas; Pontifical Catholic University of Minas Gerais; School of Economics and Business, Pontifical Catholic University of Campinas(鲍登学院; 坎皮纳斯天主教大学; 米纳斯吉拉斯天主教大学; 坎皮纳斯天主教大学经济与商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以巴西圣保罗为对象,利用谷歌地图兴趣点(POIs),通过PCA、NMF及梯度提升回归等方法,构建了低成本的亚市政级收入估算模型,其保留数据R²达0.65,可补充传统收入统计。
AI 中文摘要
中低收入国家的社会政策需要准确、最新的亚市政级收入数据,但巴西的收入数据依赖成本高昂的十年一次人口普查,其普查间隔缺口最近已超过十年。我们测试众包谷歌地图兴趣点(POIs)的构成是否可作为圣保罗市26625个普查区家庭收入的高频、低成本代理。使用从谷歌 Places 获取的一组理论驱动的 POI 类别,我们用 POI 数量表示每个普查区,通过主成分分析(PCA)和非负矩阵分解(NMF)分解这些高维稀疏特征,并训练一系列回归模型来预测普查得出的收入。在考虑数据泄漏的空间验证设计下,最佳模型(带梯度提升的 NMF)在保留数据上的 R² 达到0.65,且在特征提取方法间表现稳定。可解释的分解揭示了哪些 POI 类型承载收入信号。这些结果表明,商业众包地理空间数据可在普查间隔期间补充传统收入统计,我们还讨论了向多维贫困和能力框架扩展的可能性。
英文摘要
Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade. We test whether the composition of crowd-sourced Google Maps Points of Interest (POIs) can serve as a high-frequency, low-cost proxy for household income across the 26,625 census sectors of the municipality of Sao Paulo. Using a theoretically motivated set of POI categories retrieved from Google Places, we represent each sector by its POI counts, decompose these high-dimensional, sparse features with principal component analysis (PCA) and non-negative matrix factorization (NMF), and train a sweep of regression models to predict census-derived income. Under a data leakage-aware spatial validation design the best model (NMF with gradient boosting) attains a held-out R^2 of 0.65, with performance stable across feature-extraction methods. Interpretable decompositions reveal which POI types carry the income signal. These results suggest that commercial, crowd-sourced geospatial data can complement conventional income statistics during intercensal periods, and we discuss extensions toward multidimensional poverty and the capabilities framework.