几何自适应机制用于私有合成数据
Geometry-Adaptive Mechanisms for Private Synthetic Data
- Sharif University of Technology(谢里夫理工大学)
- University of Birmingham(伯明翰大学)
- The Alan Turing Institute(艾伦·图灵研究所)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对高维差分隐私合成数据生成,提出Adaptive Pruned-PMM机制,利用多尺度打包增长维度自适应几何,实现期望1-Wasserstein误差阶(εn)^{-1/k},并证明下界,提升低维支持场景的效用。
AI中文摘要:
在高维空间中,生成具有有意义Wasserstein效用保证的差分隐私合成数据是具有挑战性的。对于大小为\\(n\\)、定义在\\([0,1]^d\\)(其中\\(d\ge2\\))上的数据集,现有的纯\\(\varepsilon\\)-差分隐私机制实现了期望的\\(1\\)-Wasserstein误差阶为\\((\varepsilon n)^{-1/d}\\),这反映了维数灾难。虽然该速率在最坏情况下是最优的,但当数据支持在低维集合上时,它可能过于悲观。我们通过一个多尺度打包增长维度\\(k\\)来形式化这一点,该维度通过跨尺度的打包数增长来捕捉支持的几何复杂度。我们提出了\emph{Adaptive Pruned-PMM},这是一种纯\\(\varepsilon\\)-差分隐私机制,结合了私有深度选择与我们提出的He等人(2023)的Private Measure Mechanism (PMM)的剪枝变体。该机制支持更深、几何自适应的层次结构,期望运行时间为\\(O\\!\left(d(n+d)\log(\varepsilon n)\right)\\),在固定维度和隐私预算下,关于\\(n\\)是近线性的。在维度为\\(k\\)的外部多尺度打包增长条件下,我们证明,对于固定的正隐私预算和固定几何,随着\\(n\\)增长,期望的\\(1\\)-Wasserstein误差阶为\\((\varepsilon n)^{-1/k}\\)(对于\\(k>1\\))。我们还证明了在相应的内部打包增长条件下的下界,表明指数\\(1/k\\)在此框架内是尖锐的。
英文摘要:
Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.