小区域估计中的差分隐私保证
Differential Privacy Guarantees in Small Area Estimation
浏览论文内容
中文总结 AI 辅助
本研究证明贝叶斯 Fay-Herriot 模型的后验发布在固定方差分量时无需加噪即可满足 Rényi 与零集中差分隐私,并推导敏感度公式,应用于贫困与吸烟率估计。
中文摘要 AI 辅助
统计机构日益依赖小区域估计来为样本量有限的子群体生成可靠的估计值。这些估计值基于个体调查回答构建,因此机构必须确保发布这些估计值不会泄露任何单个受访者的信息。我们证明,当发布贝叶斯 Fay-Herriot 模型后验分布的单次抽样时,纯 ε-差分隐私无法实现,但若将方差分量视为固定值,该发布在无需添加任何噪声的情况下,满足 Rényi 差分隐私和零集中差分隐私下的正式隐私保证。关键见解在于后验抽样等于后验均值加上方差等于后验方差的高斯噪声。因此,该保证由直接调查估计的敏感度和后验方差决定,并且同样适用于以该噪声量添加噪声的后验均值发布。对于使用 Hájek 估计器估计的二元结果,敏感度等于该区域内最大调查权重除以权重之和。对于仅含截距的模型,我们推导出精确系数,描述一条记录的变化如何传播到每个区域的后验均值,从而给出有限样本的逐区域保证,以及同时发布所有区域的联合保证,在我们的应用中,该联合保证最多比最大逐区域保证高出几个百分点。两个应用——美国社区调查中 2,462 个公共使用微观数据区域的贫困发生率,以及华盛顿州行为风险因素监测系统中 52 个子层的吸烟率——表明,该保证受调查权重不平等性的影响远大于样本量,并且模型的收缩效应显著增强了该保证。
英文摘要
Statistical agencies increasingly rely on small area estimation to produce reliable estimates for subpopulations with limited sample sizes. These estimates are built from individual survey responses, so agencies must ensure that releasing them does not reveal information about any single respondent. We show that when a single draw from the posterior distribution of the Bayesian Fay-Herriot model is released, pure $\varepsilon$-differential privacy is unattainable, but the release satisfies formal privacy guarantees under Rényi differential privacy and zero-concentrated differential privacy without any noise being added, provided we treat the variance components as fixed. The key insight is that the posterior draw equals the posterior mean plus the Gaussian noise whose variance equals the posterior variance. The guarantee is thus governed by the sensitivity of the direct survey estimate and the posterior variance, and applies equally to a release of the posterior mean with that amount of noise added. For binary outcomes estimated with the Hájek estimator, the sensitivity equals the largest survey weight in the area divided by the sum of the weights. For the intercept-only model we derive exact coefficients describing how a change in one record propagates to every area's posterior mean, giving finite-sample per-area guarantees and a joint guarantee for releasing all areas at once that exceeds the largest per-area guarantee by at most a few percent in our applications. Two applications, poverty prevalence across 2,462 Public Use Microdata Areas in the American Community Survey and smoking prevalence across 52 substrata in the Washington state Behavioral Risk Factor Surveillance System, show that the guarantee is driven far more by the inequality of the survey weights than by the sample size, and that the shrinkage of the model tightens it substantially.
发表机构
- Washington State University(华盛顿州立大学)
- University of Maryland(马里兰大学)
- Institute for Employment Research(就业研究所)
- Ludwig-Maximilians-Universität München(慕尼黑路德维希-马克西米利安大学)
机构由 AI 辅助整理,请以论文原文为准。