描绘ε:追求图像的身份级隐私保障
Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images
浏览论文内容
中文总结 AI 辅助
本研究对比四种审计人脸生成器身份级差分隐私的方法,应用于FaceFusion和InstantID发现其存在显著身份可区分性,各方法ε估计值差异显著,后续将评估其在部分隐私机制上的权衡。
中文摘要 AI 辅助
图像到图像的人脸生成器被广泛使用,其输出与源图像之间的视觉差异有时被视为隐私的证据。审计这些系统是否满足正式的身份级(ε, δ)-差分隐私,需要在将嵌入空间观测值转换为差分隐私参数ε的估计值或界限的几种不同途径中进行选择。我们对四种适用于预训练黑盒人脸生成器的此类审计方法进行了比较研究:基于每身份敏感度的高斯机制解读(GaussMech);通过基本组合聚合的每维度核密度对数似然比(KDE-LR);通过最大均值差异经总变差距离推导得到的纯差分隐私ε的分析性总体级下界(MMD-TV);以及对交叉验证分类器的袋外ROC进行的假设检验评估(ROC-HT)。对于每种方法,我们明确了其假设、超参数依赖性、有限样本限制以及其ε估计值具有信息性的适用场景。将这些审计方法应用于FaceFusion和InstantID,跨多个身份编码器和参考数据集,结果一致显示存在显著的身份可区分性,同时报告的ε估计值却明显不同,反映了每种方法的不同假设和有限样本处理方式。在这一高可区分性场景下,实验不支持对四种方法进行可靠排序,它们的相对权衡应在部分隐私机制上进行评估,我们将此确定为后续自然研究方向。所得框架将这些审计方法置于统一的身份级审计场景中,并阐明了它们的假设和有限样本处理如何塑造最终的差分隐私估计值。
英文摘要
Several methods for auditing privacy in embedding spaces report a number called "epsilon." The common name is misleading: one number may be a heuristic score, another may come from a valid population inequality but ignore sampling uncertainty, and a third may be a confidence bound. This paper asks when such a number is evidence about differential privacy. We study four approaches based on Gaussian calibration, marginal kernel-density ratios, maximum mean discrepancy (MMD), and classifier hypothesis tests. The first two are modeling diagnostics. The MMD and classifier approaches use valid population lower bounds, but only the classifier approach is given a separate test set and a finite-sample confidence calculation. To see how these distinctions matter, we build a synthetic benchmark in which the true privacy value is known. It includes pure-DP Laplace mechanisms, approximate-DP Gaussian mechanisms, an exact privacy null, two identity geometries, and three sample sizes. At the null, the Gaussian and kernel-density diagnostics remain large. A direct conversion of an empirical ROC curve often returns infinity even though the AUC is near chance and no threshold separates the samples perfectly. The MMD value increases as the distributions become easier to distinguish, but the sample estimate is small relative to the known reference and is not a lower confidence bound. By contrast, a classifier chosen on development data and evaluated on untouched test data gives simultaneous lower confidence bounds under the stated iid model. We also apply the methods to FaceFusion and InstantID. That case study shows how the methods behave on face data but does not calibrate them because the generators have no known privacy value. The main lesson is that an epsilon-like number is meaningful only together with its assumptions and its finite-sample interpretation.
发表机构
- Kansas State University(堪萨斯州立大学)
- Louisiana State University(路易斯安那州立大学)
机构由 AI 辅助整理,请以论文原文为准。