arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于比较人口普查数据收集方法的总统计误差框架

A Total Statistical Error Framework for Comparing Census Data Collection Methods

Siu-MIng Tam, Anders Holmberg

arXiv 2608.00511首次发表:更新:

AI 中文总结

该研究构建了总统计误差框架,区分两种评估范式,利用2021年澳大利亚普查数据对比k近邻与随机森林插补,发现加入普查无应答指标可大幅校正依赖型无应答导致的偏差,为人口普查数据收集方法比较提供了工具。

AI 中文摘要

人口普查越来越依赖插补来为无应答住户分配常住地址,但目前尚无正式统计框架可比较不同插补方法在覆盖范围与地址准确性综合表现上的优劣。我们在总统计误差(Total Statistical Error, TSE)范式下构建了该框架,该框架不仅限于基于调查的数据收集,同样适用于基于登记和行政数据源的情况。核心指标是个体层面的二元正确性指示器:当且仅当某人被计入普查且被分配到正确的常住地址时,该指示器取值为1;其补集即为个体总统计误差,将其在全人口层面汇总得到的正确率可作为直接比较不同方法的依据。我们区分了两种评估范式:范式I中,普查后调查(Post-Enumeration Survey, PES)为概率样本提供参考值;范式II中,完整基准可直接计算正确率。我们利用2021年澳大利亚人口普查微观数据对该框架进行了演示,在非随机缺失(Not-Missing-At-Random, NMAR)机制下比较了k近邻(k-nearest-neighbour)和随机森林(random forest)插补方法。一项关键发现是,当普查与模拟PES之间存在依赖型无应答且亚组删除率较小时,简单的范式I亚组估计会出现严重偏差。在PES响应倾向模型中加入普查无应答指标可将该偏差降低约90%;模拟实验表明,偏差校正主要由无应答指标单独驱动,额外加入的插补集A值贡献的校正效果可忽略不计。当担忧普查与PES之间存在依赖型无应答时,该策略为首选方案。

英文摘要

Population censuses increasingly rely on imputation to assign usual-residence addresses for non-responding dwellings, yet no formal statistical framework has existed for comparing competing imputation methods on their combined coverage and address-accuracy performance. We develop such a framework within a Total Statistical Error (TSE) paradigm that is not restricted to survey-based data collection and applies equally to register-based and administrative data sources. The central quantity is a unit-level binary correctness indicator, equal to one if and only if a person is both enumerated and assigned to the correct usual-residence address; its complement is the unit TSE, and aggregating over the population yields a correctness rate that serves as the basis for head-to-head method comparison. We distinguish two assessment paradigms: Paradigm I, in which a post-enumeration survey (PES) provides reference values for a probability sample, and Paradigm II, in which a complete benchmark makes the correctness rate directly computable. We illustrate the framework using 2021 Australian census microdata, comparing k-nearest-neighbour and random forest imputation under a not-missing-at-random (NMAR) mechanism. A key finding is that dependent nonresponse between the census and a simulated PES causes naive Paradigm I subgroup estimates to be severely biased when subgroup deletion rates are small. Augmenting the PES response propensity model with the census nonresponse indicator reduces this bias by approximately 90%; simulation experiments show that the nonresponse indicator alone drives the correction, with the additionally included imputed set A values contributing negligible further reduction. This is the preferred strategy whenever dependent nonresponse between the census and PES is a concern.

Comments24 pages, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑