AI 中文总结
该研究针对局部差分隐私下的实值数据,提出拉普拉斯噪声下非参数极大似然估计的有限维重构形式,分析其收敛性与隐私噪声的关联,明确了NPMLE保持一致性的噪声速率条件。
AI 中文摘要
局部差分隐私(LDP)通过在发布前对每个测量值进行扰动来保护数据集中的个体。对于实值数据,一种广泛使用的机制是添加拉普拉斯噪声。我们研究在独立同分布采样模型下,通过非参数极大似然估计(NPMLE)从私有化观测值估计潜在机密数据分布的问题。我们首先证明,在拉普拉斯卷积模型下,NPMLE具有有限维重构形式,其支撑集被限制在观测值集合上。这将原有的对所有混合分布的无限维优化,简化为对混合权重的n维凸优化,其中n为样本量。随后我们研究NPMLE在1- Wasserstein距离下的统计收敛性,并明确将其收敛速率与隐私噪声尺度关联起来。考虑隐私噪声水平随样本量变化的情况,我们的分析表明,当拉普拉斯噪声以慢于n^(3/16)的速率增长时,NPMLE仍保持一致性;相反,当拉普拉斯噪声为√n量级或更大时,不存在任何估计器能实现潜在分布的一致恢复。
英文摘要
Local differential privacy (LDP) protects individuals in a dataset by perturbing each measurement before release. For real-valued data, a widely used mechanism is additive Laplace noise. We study the problem of estimating the distribution of the latent confidential data from the privatized observations via the nonparametric maximum likelihood estimator (NPMLE) under an i.i.d. sampling model. We first show that under the Laplace convolution model, the NPMLE admits a finite-dimensional reformulation in which the support is restricted to the observation set. This reduces the original infinite-dimensional optimization over all mixing distributions to an $n$-dimensional convex optimization over mixture weights, where $n$ is the sample size. We then study the statistical convergence of the NPMLE under the 1-Wasserstein distance, and explicitly connect its convergence rate with the privacy noise scale. Allowing the privacy noise level to change with the sample size, our analysis shows that the NPMLE remains consistent when the Laplace noise grows at a rate slower than $n^{3/16}$. Conversely, when the Laplace noise is of the order $\sqrt n$ or larger, no estimator can achieve uniformly consistent recovery of the latent distribution.
Comments32 pages, 4 figures