发表机构
Aalborg University; Aalborg University Hospital(奥尔堡大学; 奥尔堡大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于正态-逆Wishart先验的多元正态数据贝叶斯合成器,揭示合成数据泄露风险与原始数据集规模、特征维度及合成信息量的关系,并将其应用于多基因风险评分数据集生成。
AI 中文摘要
贝叶斯合成是一种通过从后验预测分布中采样生成合成数据的流行方法,用于对敏感个人数据进行隐私保护。然而,属性泄露风险如何受特征维度、原始数据集的个体数量以及所发布合成信息量的影响,目前尚不清楚。我们针对多元正态数据提出了一种数学上可处理的贝叶斯合成器,其基于均值向量和协方差矩阵的共轭正态-逆Wishart(NIW)先验。这种共轭结构可产生闭式后验分布,使我们能够直接研究在不同数据和发布设置下对手推断记录的能力。随后,我们通过多次模拟展示了合成数据生成的若干直观特性:具体而言,我们发现泄露风险随原始数据集规模增大而降低,但随特征空间维度及所发布合成信息量的增加而升高,无论这种合成信息量的增加是通过发布更大的合成数据集还是多个生成器实现。最后,所提出的合成器被用于生成多基因风险评分数据集的合成版本,该合成数据展现出与原始数据相当的分布特性。
英文摘要
Bayesian synthesis, which generates synthetic data by sampling from the posterior predictive distribution, is a popular approach for privatizing sensitive personal data. However, how attribute disclosure risk is affected by feature dimensionality, the number of individuals in the original dataset, and the amount of released synthetic information remains poorly understood. We propose a mathematically tractable Bayesian synthesizer for multivariate normal data based on a conjugate normal-inverse-Wishart prior for the mean vector and covariance matrix. The conjugate structure yields closed-form posteriors and enables direct investigation of an adversary's ability to infer records under different data and release settings. We then demonstrate several intuitive properties of synthetic data generation through several simulations. Specifically, we show that disclosure risk decreases with the size of the original dataset, but increases with the dimensionality of the feature space and the amount of synthetic information released, whether through the release of larger synthetic datasets or multiple generator realizations. Finally, the proposed synthesizer was used to generate synthetic versions of a polygenic risk score dataset, with the synthetic data exhibiting distributional properties comparable to those of the original data.