arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

私有合成测度的内蕴维数Wasserstein保证

Intrinsic-Dimensional Wasserstein Guarantees for Private Synthetic Measures

Yiyun He

arXiv 2609.17624首次发表:更新:

发表机构

University of California, San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于PrivTree的差分隐私合成测度方法,其Wasserstein误差由内蕴维数而非环境维数决定,并达到极小极大最优率,同时引入平移技术消除常数对维数的指数依赖。

AI 中文摘要

我们研究$[0,1]^d$中$n$个点的$\varepsilon$-差分隐私合成测度,通过应用现有的PrivTree算法构造自适应二叉划分,然后私有地发布其叶质量。我们考虑最坏情况数据模型,不假设任何采样或总体分布。对于$d\ge2$,合成测度的1-Wasserstein误差为$\widetilde O_d((\varepsilon n)^{-1/d})$,与极小极大下界相比达到最优(至多相差对数因子)。此外,对于$d\ge3$且$2<s\le d$,如果数据集在相关有限尺度范围$r$上的覆盖数至多为$Ar^{-s}$,则期望误差改进为$\widetilde O_{d,s}((\varepsilon n)^{-1/s})$。因此,速率取决于有限尺度的内蕴维数而非环境维数,且无需恢复低维流形。我们还引入一种平移技术,以进一步避免常数对环境维数$d$的指数依赖。

英文摘要

We study an $\varepsilon$-differentially private synthetic measure for $n$ points in $[0,1]^d$ by applying the existing PrivTree algorithm to construct an adaptive binary partition and then privately releasing its leaf masses. We consider the worst-case data model without any sampling or population-distribution assumption. The 1-Wasserstein error of the synthetic measure is $\widetilde O_d((\varepsilon n)^{-1/d})$ for $d\ge2$, which is optimal compared to the minimax lower bound up to a logarithmic factor. Moreover, for $d\ge3$ and $2<s\le d$, if the data set has covering number at most $Ar^{-s}$ over the relevant finite range of scales $r$, the expected error improves to $\widetilde O_{d,s}((\varepsilon n)^{-1/s})$. Thus the rate depends on a finite-scale intrinsic dimension rather than the ambient dimension, without requiring the recovery of a low-dimensional manifold. We also introduce a shifting technique to further avoid the exponential dependence of the constant on the ambient dimension $d$.

Comments30 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑