AI 中文总结
研究在存在线性不变量(如数据库特定聚合查询受限)情况下的差分隐私实现问题,提出高熵差分隐私实现方法,能维持聚合不变量并给出隐私保证,其理论证明有新贡献且方法在正态混合模型采样中有广泛用途。
AI 中文摘要
差分隐私是确保数据隐私的标准,广泛应用于主要数据发布中。常见的差分隐私实现是向数据库添加独立的高斯或拉普拉斯噪声。但对于数据库的聚合(线性)查询可能被排除在隐私预算之外。在聚合约束下,噪声向量不再独立,传统差分隐私保证需重新评估。我们提出一种高熵差分隐私实现,以概率 1 或指数接近 1 的概率维持聚合不变量,并推导不变量下的隐私保证。理论证明涵盖了相关矩阵零空间一个开放问题的部分解决方案。此外,该方法在满足线性等式约束的正态混合模型采样中具有广泛应用。
英文摘要
Differential privacy is the standard for ensuring data privacy and is widely used in major data publications, including reporting results from the U.S. decennial census. Common implementation of differential privacy uses independent Gaussian or Laplace noise addition to the database. However, there could be aggregate (linear) queries to the database that are excluded from the privacy budget, for example, state totals that can not be perturbed due to constitutional mandates. Any implementation of a differential privacy is required to honor these constraints, also referred to as invariants. Under aggregation constraints, the noise vector is no longer independent and the traditional differential privacy guarantees have to be re-evaluated. We propose a high entropy differential privacy implementation that maintains the aggregation invariants with probability one or exponentially close to one and derive the privacy guarantees for the implementation under the invariants. The theoretical proof covers a partial solution to an open question about the null space of correlation matrices. Moreover, the methodology has general use in the context of sampling from normal mixture models under linear equality constraints.