通过PAC-私有适配扩散模型保护图像合成中的敏感数据
Protecting Sensitive Data in Image Synthesis via PAC-Private Adaptation for Diffusion Models
浏览论文内容
中文总结 AI 辅助
针对扩散模型合成数据易受重建攻击且DP方法效用损失大的问题,提出PAC-私有适配方法,通过LoRA或文本反演学习紧凑组件并校准各向异性噪声,实现直接重建隐私,在少样本个性化和全数据集合成中优于DP。
中文摘要 AI 辅助
合成数据越来越多地被用作共享敏感记录的替代方案。然而,合成数据生成并不能保证隐私,因为在敏感数据上训练或适配的扩散模型仍然容易受到重建攻击。此外,虽然使用差分隐私(DP)的方法(如DP-SGD)能够实现可证明私有的扩散模型训练,但它们所需的重复梯度裁剪和噪声注入会导致显著的效用损失。基于DP的隐私的一个重要局限性在于,尽管它与重建隐私(RP)具有可证明的关系,但这种关系是间接的。RP的定义是限制对手对敏感数据的后验分布与先验分布之间的差异,而DP则通过限制输出对单个记录变化的敏感性来提供保证。这种间接性是效用损失的一个重要来源。为了解决这一问题,我们提出了一种PAC-私有扩散模型适配方法以实现重建隐私。由于PAC-隐私直接针对相对于先验的后验优势进行定义,因此它直接涉及RP。为了在高维空间中实现可扩展的PAC私有化,我们首先使用LoRA或文本反演学习一个紧凑的数据相关扩散模型组件,然后根据重复机制输出的协方差校准各向异性高斯噪声。与DP-SGD不同,我们的方法在优化后仅扰动一次所学组件,从而避免了跨梯度更新的隐私组合。我们在少样本概念个性化和全数据集图像合成上评估了该框架,并表明所提出的方法在实现相同重建隐私的同时,比DP更好地保留了主体身份、生成质量和下游分类准确性。
英文摘要
Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover, while approaches that use differential privacy (DP), such as DP-SGD, achieve provably private diffusion model training, the repeated gradient clipping and noise injection they require result in significant utility loss. An important limitation of DP-based privacy is that, although it has a provable relationship to reconstruction privacy (RP), that relationship is indirect. RP is defined in terms of limiting how much an adversary's posterior distribution over sensitive data differs from the prior, whereas DP provides guarantees by bounding the sensitivity of outputs to changes in individual records. This indirection is an important source of the utility loss. To address this, we propose a PAC-private diffusion model adaptation to achieve reconstruction privacy. Since PAC-privacy is defined directly with respect to posterior advantage over the prior, it directly implicates RP. To obtain scalable PAC privatization in high dimensions, we first learn a compact data-dependent diffusion model component using LoRA or Textual Inversion, and then calibrate anisotropic Gaussian noise from the covariance of repeated mechanism outputs. Unlike DP-SGD, our method perturbs the learned component only once after optimization, thereby avoiding privacy composition across gradient updates. We evaluate the framework on few-shot concept personalization and full-dataset image synthesis, and show that the proposed approach better preserves subject identity, generation quality, and downstream classification accuracy than DP while achieving the same reconstruction privacy.
发表机构
- Washington University in St. Louis(圣路易斯华盛顿大学)
- Virginia Tech(弗吉尼亚理工大学)
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。