一种用于普查 tract 人口估算的混合状态空间方法
A Hybrid State-Space Approach for Census-Tract Population Estimation
- Rutgers University(罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出 MambaPop 方法,将行政单元卫星图像作为序列建模输入,直接估算人口,消除分解步骤,在 8.4 万个美国本土普查 tract 上达到与 YOLOv11 相当的 MAE。
AI中文摘要:
序列模型是大型语言模型及越来越多的先进图像识别技术背后的架构家族,它重新定义了机器从高维数据中学习的方式。然而,作为基础设施规划、公共卫生和灾害应对基础的卫星图像人口估算任务,却几乎未从中受益:主流系统仍将人口绑定到统一栅格,通过辅助数据(如 WorldPop 和 LandScan)构建的权重表面将普查数据分解到网格单元,这可能引入系统性空间偏差,同时用卷积神经网络预测每个网格单元的人口。这种方法丢弃了收集普查数据时实际采用的行政单元结构。我们通过 MambaPop 填补这一空白,该模型将每个行政单元呈现为单一的多边形掩码卫星图像,并将 tract 级人口估算视为对其图像块的序列建模问题,直接将每个 tract 图像与其人口标签配对,完全消除了分解步骤。MambaPop 基于混合状态空间-注意力 MambaVision 骨干网络构建,据我们所知,它是首个直接从行政单元自身图像学习人口的方法,也是首个将基于状态空间的(Mamba)混合架构应用于人口估算任务的方法。在 2020 年人口普查的约 84000 个美国本土普查 tract 中,MambaPop 达到的平均绝对误差(MAE)为每个 tract 1141 人,与最强的卷积基线(YOLOv11,MAE 1122)相当。
英文摘要:
Sequence models---the architecture family behind large language models and, increasingly, state-of-the-art image recognition---have redefined how machines learn from high-dimensional data. Yet population estimation from satellite imagery, a task that underpins infrastructure planning, public health, and disaster response, has scarcely benefited: leading systems still bind population to a uniform raster, disaggregating census counts onto grid cells through weighting surfaces built from ancillary data (e.g., in WorldPop and LandScan), which can introduce systematic spatial bias, and predicting population per grid cell with convolutional neural networks. In this approach, the administrative-unit structure in which the census was actually collected is discarded. We close this gap with MambaPop, which renders each administrative unit as a single polygon-masked satellite image and treats tract-level population estimation as a sequence-modeling problem over its image patches, pairing each tract image directly with its population label and eliminating the disaggregation step entirely. Built on the hybrid state-space--attention MambaVision backbone, MambaPop is, to our knowledge, the first method to learn population directly from an administrative unit's own image as well as the first to apply a state-space based (Mamba) hybrid architecture to the population estimation task. Across all $\sim$84{,}000 contiguous-US census tracts of the 2020 census, MambaPop attains a mean absolute error (MAE) of $1{,}141$ persons per tract, matching the strongest convolutional baseline (YOLOv11, MAE $1{,}122$).