arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向策略感知的私有数据合成的软投票方法

Soft Voting for Policy-Aware Private Data Synthesis

Yingge Hu, Gautham Ramesh Babu, Mostafa Milani

arXiv 2610.11285首次发表:更新:

发表机构

Western University(韦仕敦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Private Evolution等DP合成器的硬投票在策略图下无法降噪声的问题,提出BF-Soft软投票方法,在强隐私预算下可降低误差,公共数据试点还能预测软投票的适用性。

AI 中文摘要

Blowfish隐私通过仅保护数据所有者指定为策略图边的属性值替换,放松了差分隐私(DP)的要求。更稀疏的策略可降低机制所需的噪声,但前提是在受保护替换下的发布统计量变化小于任意DP邻居间的变化。我们针对进化型、近邻DP合成器(如Private Evolution(PE)及其表格实例Tab-PE)研究该问题,这类合成器会对私有记录与候选群体打分并发布带噪投票直方图。它们的硬投票在每个候选的决策区域内为常数,在边界处跳变,因此只要至少一个受保护替换跨越边界,无论该替换多短,其策略特定敏感度都等于完整最坏情况值。由于我们考察的每一轮都存在此类替换,策略图未实现噪声降低。我们提出BF-Soft,一种温度平滑的软投票,其响应随距离逐渐变化,敏感度在策略图的可达性和温度上有紧密闭式界,与候选数量无关,且该界可在合成前计算一次。它还能仅从策略预测策略感知平滑无法大幅降低噪声的情况:保护平坦分类或二元属性会使可达性达到最大值。在窄数值策略下的真实和合成数据集上,BF-Soft在强隐私预算下相比硬投票降低了误差,而在弱预算下优势反转。一项公共数据试点可在不消耗私有预算的情况下预测软投票何时有益。

英文摘要

Blowfish privacy relaxes differential privacy (DP) by protecting only the attribute-value substitutions a data owner specifies as edges of a policy graph. A sparser policy can reduce the noise required by a mechanism, but only when the released statistic changes less across protected substitutions than across arbitrary DP neighbors. We study this question for evolutionary, nearest-neighbor DP synthesizers such as Private Evolution (PE) and its tabular instantiation Tab-PE, which score private records against a candidate population and release a noisy vote histogram. Their hard vote is constant inside each candidate's decision region and jumps at its boundary. Its policy-specific sensitivity therefore equals the full worst-case value whenever at least one protected substitution crosses a boundary, regardless of how short that substitution is. Because every round we examined contained such a substitution, the policy graph gave no reduction in noise. We propose BF-Soft, a temperature-smoothed soft vote whose response changes gradually with distance. Its sensitivity has a tight closed-form bound in the policy graph's reach and the temperature, independent of the number of candidates, and the bound can be computed once before synthesis. It also predicts from the policy alone when policy-aware smoothing cannot substantially reduce noise: protecting a flat categorical or binary attribute drives the reach to its maximum. On real and synthetic datasets under narrow numeric policies, BF-Soft reduces error relative to hard voting at strong privacy budgets, while the advantage reverses at weaker budgets. A public-data pilot predicts when soft voting is beneficial without spending private budget.

Comments22 pages, 11 figures, 4 tables. Extended version with full proofs and additional experiments. Code: https://github.com/ayanamirei629/Soft-Voting-for-Policy-Aware-Private-Data-Synthesis

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑