均匀有界分布下的收缩性与隐私统计推断
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions
AI总结:
本文提出$c$-内部逐点最大泄漏(PML)作为局部差分隐私的推广,用于更灵活的收缩分析,并在二元假设检验和均值估计中实现渐近最优的隐私推断,且在高规则性下不增加样本复杂度。
AI中文摘要:
我们研究 $c$-内部逐点最大泄漏(PML)作为收缩分析和披露控制的工具。基于最大泄漏的强对抗威胁模型,$c$-内部PML将局部差分隐私(LDP)推广到密度被常数 $c>0$ 均匀有界远离零的数据生成分布。将 $c$-内部PML视为核上的代数约束,可以得到比标准LDP更灵活(且通常更紧)的收缩分析。我们提供了Dobrushin系数的紧界,并界定了Hockeystick散度的收缩系数。我们进一步在$c$-内部PML约束下,当散度的输入分布被限制在$c$-内部时,推导了$f$-散度的强数据处理不等式。这些结果超越了纯LDP的范围,覆盖了更大的核类,例如任意随机矩阵。我们将结果应用于极小极大理论,并在$c$-内部PML约束下为二元假设检验和均值估计提供了渐近最优策略。结果表明,使用PML的披露控制允许分析者以更细致的方式推理系统:例如,它使我们能够量化确定性系统的隐私泄漏,并可以针对任意分布假设给出精确的对抗性保证。有趣的是,披露分析中的一个反复出现的主题是,如果隐私问题相对规则(即密度界$c$较大),则可以在不产生额外样本复杂度成本的情况下进行私有推断。
英文摘要:
We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.