发表机构
The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出差分隐私非参数模态学习方法DP-GRAMS,实现多峰分布密度模态的高概率恢复,扩展出DP-PMS回归与DP-GRAMS-C聚类方法,在合成及真实数据上展现良好隐私-效用权衡。
AI 中文摘要
密度模态为多峰分布提供了局部且可解释的摘要,但在严格差分隐私约束下对其进行估计的研究仍十分有限。我们研究了在局部光滑性、曲率及分离条件下,多元分布密度模态的差分隐私恢复问题。提出了DP-GRAMS这一受均值漂移启发的方法,该方法对差分隐私得分估计器执行带噪声的上升操作。假设密度局部属于光滑参数β>2的Hölder类,我们的得分估计器使用降低偏差的高阶核,随后通过梯度裁剪和校准高斯噪声在梯度上升步骤中实现隐私保护。一种私有初始化方案结合了密度感知效用与抑制规则,当k≍Mlogn、在公共h_DAP网格上采样且抑制半径ρ_init≍(logn)^{-1/d}时,通过连续抑制竞争区域内选定的局部邻域,实现了模态盆地的高概率覆盖,而多次启动产生的相关噪声支持在单一(ε,δ)-差分隐私保证下联合发布。我们证明所有总体模态均能以高概率被恢复,并建立了渐近误差率形式为O((logn/n)^(2(β-1)/(d+2β))) + O((polylog(n,δ)/(n²ε²))^((β-1)/(d+β)))。我们还给出了私有模态估计的极小极大下界,并证明我们的估计器在均方误差上仅差一个对数因子,接近最优。我们提供了两种自然扩展:私有模态回归方法DP-PMS,以及聚类流程DP-GRAMS-C。在合成数据与真实数据上的大量实验表明,相较于常见基线,我们的方法具有良好的隐私-效用权衡。
英文摘要
Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. We study differentially private recovery of density modes for multivariate distributions under local smoothness, curvature, and separation conditions. We propose DP-GRAMS, a mean-shift inspired method that performs noisy ascent on a differentially private score estimator. Assuming the density belongs locally to a Hölder class with smoothness parameter $β> 2$, our score estimator uses bias-reducing higher-order kernels, and then enforces privacy in the gradient ascent steps via gradient clipping and calibrated Gaussian noise. A private initialization scheme combines a density-aware utility with a suppression rule and, with $k\asymp M\log n$ draws over a public $h_{\mathrm{DAP}}$-grid and suppression radius $ρ_{\mathrm{init}}\asymp (\log n)^{-1/d}$, achieves high-probability coverage of the modal basins by successively suppressing selected local neighborhoods in competitive regions, while correlated noise across multiple starts enables joint release under a single $(\varepsilon,δ)$-differential privacy guarantee. We prove that all population modes are recovered with high probability and establish asymptotic error rates of the form $O\!\left((\tfrac{\log n}{n})^{\frac{2(β-1)}{d+2β}}\right) + O\!\left((\tfrac{\mathrm{polylog}(n,δ)}{n^2\varepsilon^2})^{\frac{β-1}{d+β}}\right)$. We also provide minimax lower bounds for private mode estimation, and show that our estimators are nearly optimal, up to a logarithmic factor in the MSE. We present two natural extensions: DP-PMS, a private modal-regression method, and DP-GRAMS-C, a clustering pipeline. Extensive experiments on synthetic and real data demonstrate favorable privacy-utility trade-offs relative to common baselines.