arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分块的私有生成自助法

Private Generative Bootstrap via Blocking

Jinwon Sohn, Veronika Ročková

arXiv 2608.02480首次发表:更新:

发表机构

Booth School of Business, University of Chicago(芝加哥大学布斯商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出私有生成贝叶斯自助法(PGBB),采用分块策略结合差分隐私与摊销推理,在模拟及实际数据应用中实现了更优的私有不确定性量化。

AI 中文摘要

随着AI系统获取个体信息的途径日益增多,在报告统计结果时保护隐私变得尤为重要,而对这类结果的不确定性报告进行私有化处理也同样关键。为此,我们采用了贝叶斯无似然框架,并对后验分布的模拟进行私有化处理。特别地,我们提出了一种新的私有贝叶斯自助法实例,该方法采用分块策略,不再为每个个体分配独特的随机权重,而是将个体随机分组并为每个组分配单一权重。通过将个体的贡献隐藏在组内,我们强化了差分隐私机制。我们利用摊销推理将私有学习与后验采样解耦,通过在训练过程中添加校准后的噪声,以私有方式学习从观测权重到后验样本的前向映射,后续的后验抽样无需额外的隐私和计算预算。我们将该方法命名为私有生成贝叶斯自助法(PGBB)。我们建立了差分隐私保证,分析了其向非私有分块自助法目标的收敛性,量化了普通贝叶斯自助法后验与分块贝叶斯自助法后验之间的差异。此外,我们推导了分块狄利克雷浓度参数的无数据调优方法,该方法可渐近恢复后验离散度。我们还表明,单次拟合PGBB可同时支持一系列基于损失的决策规则,且无需额外的隐私成本。在模拟实验以及对美国人口普查教育回报数据、美国新生儿体重分位数数据的应用中,PGBB在私有不确定性量化方面表现具有竞争力,且在常见设置下优于需要指定数据生成模型的私有贝叶斯替代方法。

英文摘要

With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers. To this end, we adopt a Bayesian likelihood-free framework and make simulation from the posterior private. In particular, we propose a new private instantiation of the Bayesian bootstrap using a blocking strategy. Rather than assigning idiosyncratic random weights to each individual, we randomly group individuals and assign a single weight to each group. By concealing individuals' contributions within a group, we fortify differential privacy gates. We harness amortized inference that decouples private learning from posterior sampling. A push-forward map from observation weights to posterior samples is learned privately by adding calibrated noise during training. Subsequent posterior draws require no additional privacy and computation budget. We call the resulting method the Private Generative Bayesian Bootstrap (PGBB). We establish a differential privacy guarantee, analyze convergence to the non-private blocked-bootstrap target, and quantify the discrepancy between the ordinary and blocked Bayesian-bootstrap posteriors. In addition, we derive data-free tuning of the block Dirichlet concentration parameter that restores posterior dispersion asymptotically. We also show a single fit of PGBB can support a family of loss-based decision rules simultaneously without additional privacy cost. In simulations and in applications to U.S. Census returns to schooling and U.S. natality birthweight quantiles, PGBB gives competitive private uncertainty quantification and improves over private Bayesian alternatives that require a specified data-generating model in common settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑