差分隐私下数据重建的尖锐转变
A Sharp Transition in Data Reconstruction under Differential Privacy
浏览论文内容
中文总结 AI 辅助
本研究揭示差分隐私下数据重建存在尖锐转变:当隐私预算远小于数据维度时重建不可能,远大于时可行,且转变取决于数据的有效维度,为隐私预算选择提供理论指导。
中文摘要 AI 辅助
数据重建攻击在经验上已成功地从学习模型中恢复训练样本,引发了隐私担忧,并促使人们提出在应对未来威胁时仍保持有效保证的防御措施。虽然差分隐私(DP)提供了形式上的保护,但选择隐私预算仍然是一个挑战:较小的预算会严重降低效用,但很难量化在不允许准确重建的情况下预算可以设置多大。在本工作中,我们研究知情的攻击者,其目标是从一个ρ-零集中差分隐私模型中重建单个d维训练样本,并已知所有其他训练数据。我们的主要贡献是建立了数据重建在ρ ≍ d处的尖锐转变:一方面,我们为任何私有机制和任何攻击推导了基于熵的下界,刻画了一组目标先验,对于这些先验,当ρ ≪ d时,重建在信息论上是不可能的;另一方面,我们分析了带有输出扰动的私有线性回归上的一种简单攻击,表明当ρ ≫ d时,重建实际上是可行的。值得注意的是,对于位于s维子空间中的数据,转变移至ρ ≍ s,这表明保证充分保护的隐私预算必须根据数据的有效维度来评估。我们通过合成数据和自然图像(CIFAR-10、ImageNet)上的实验验证了我们的发现。
英文摘要
Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
发表机构
- Institute of Science and Technology Austria(奥地利科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。