arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一隐私核算:信息等价与信息损失

Unifying Privacy Accounting: Information Equivalence and Information Loss

Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang

arXiv 2610.02414首次发表:更新:

AI 中文总结

本文统一了四种差分隐私曲线概念,证明信息等价性,并量化RDP与zCDP间的信息损失,指出保留完整RDP曲线可显著降低噪声方差并提升DP-SGD准确率。

AI 中文摘要

差分隐私(DP)存在多种概念,但选择其中之一可能会影响隐私分析和效用。在本文中,我们在一个统一的信息论框架内考虑四种主流的基于曲线的隐私概念。对于固定的输出分布有序对,我们建立了$(\varepsilon,\delta)$-DP的两个方向性隐私轮廓、假设检验权衡函数对以及扩展隐私损失分布之间的信息等价性。精确的Rényi差分隐私(RDP)曲线在某个大于一的阶数处有限时,也加入该等价类。在此温和条件下,选择这些概念仅改变其语义解释和计算要求。相反,取方向性隐私轮廓的最大值或将RDP曲线压缩为单个零集中差分隐私(zCDP)参数可能会丢失信息。我们量化了标准噪声机制下精确RDP曲线与其zCDP界之间的信息损失。对于高斯噪声,该差距为零,但对于高斯混合、拉普拉斯、离散高斯和泊松子采样高斯机制,该差距通常为正。此外,该差距随独立组合机制的数量线性增长。我们的信息论视角具有实际意义。在相同的认证隐私水平下,保留完整的RDP曲线而非使用zCDP,对于规模与美国社区调查相当的工作负载中的高斯混合噪声,可将所需噪声方差降低高达$45\\%$。对于在泊松子采样下对Fashion-MNIST进行DP-SGD,当两者校准到相同的$(\varepsilon,\delta)$保证时,基于RDP的隐私核算器相比基于zCDP的核算器,测试准确率可提高高达$8.73$个百分点。

英文摘要

Differential privacy (DP) admits several notions, but the choice among them may affect both privacy analysis and utility. In this paper, we consider four mainstream curve-based privacy notions within a unified information-theoretic framework. For a fixed ordered pair of output distributions, we establish information equivalence among the two directional privacy profiles of $(\varepsilon,δ)$-DP, the pair of hypothesis-testing trade-off functions, and the extended privacy-loss distribution. The exact Rényi differential privacy (RDP) curve joins this equivalence class whenever it is finite at some order greater than one. Under this mild condition, choosing among these notions changes only their semantic interpretation and computational requirements. In contrast, taking the maximum of the directional privacy profiles or compressing the RDP curve into a single zero-concentrated differential privacy (zCDP) parameter can lose information. We quantify the information loss between the exact RDP curve and its zCDP bound for standard noise mechanisms. This gap is zero for Gaussian noise but generally positive for Gaussian-mixture, Laplace, discrete Gaussian, and Poisson-subsampled Gaussian mechanisms. Moreover, this gap grows linearly with the number of independently composed mechanisms. Our information-theoretic perspective has practical consequences. At the same certified privacy level, retaining the full RDP curve rather than using zCDP reduces the required noise variance by up to $45\%$ for Gaussian-mixture noise in workloads comparable in size to the American Community Survey. For DP-SGD on Fashion-MNIST under Poisson subsampling, an RDP-based privacy accountant improves test accuracy by up to $8.73$ percentage points compared to a zCDP-based accountant when both are calibrated to the same $(\varepsilon,δ)$ guarantee.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑