发表机构
Bar-Ilan University(巴伊兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种逐分量私有学习框架,利用LabelBoost恢复可实现性,将学习开销从O(k/ε)降至O(√k/ε),并改进半空间与布尔组合的样本复杂度。
AI 中文摘要
我们研究可实现设定下的差分隐私学习问题,其中假设由 $k$ 个分量指定。直接迭代私有分量学习器会遇到一个简单困难:即使下一个分量在局部是准确的,对其的近似选择也可能破坏带标签样本的精确可实现性。我们使用 Beimel、Nissim 和 Stemmer [SODA '15, Algorithmica '21] 的 LabelBoost 过程恢复可实现性,并通过两个交替的储备池回收数据。对于目标隐私 $\varepsilon$,所得学习器相对于单个分量学习步骤在目标精度 $\Theta(\alpha/k)$ 下的主动样本需求,仅需支付 $\widetilde O(\sqrt{k}/\varepsilon)$ 的额外开销。对于在大小为 $L$ 的有限坐标网格上学习 $d$ 维半空间,精确可实现性使得直接分量深度目标成为拟凹函数。将该框架与 Nissim、Tsfadia 和 Yan [SODA '26] 的 IPConcave 算法以及 Cohen、Lyu、Nelson、Sarl'os 和 Stemmer [STOC '23] 的拟凹优化器实例化,得到可实现样本复杂度为 $\widetilde{O}\left(\frac{1}{\varepsilon \alpha}\cdot \min\{d^{2.5} \log^*L, \\\\:\\\\: d^{2.5} + d^{1.5} 2^{\log^*L}\}\right)$,这改进了先前已知的界 $\widetilde{O}\left(\frac{1}{\varepsilon \alpha}\cdot\min\{\frac{1}{\alpha}\cdot d^{5.5}\log^*L,\\\\:\\\\: d^{2.5}2^{\log^*L}\}\right)$。我们还将该框架应用于布尔组合:给定类别 $H_1,\ldots,H_k$ 的恰当私有学习器,我们为任意固定布尔函数 $G:\{0,1\}^k\to\{0,1\}$ 获得 $G(H_1,\ldots,H_k)$ 的恰当私有学习器。与 Alon、Beimel、Moran 和 Stemmer [COLT '20] 的闭包定理相比,这将常见分量样本界上的额外开销从 $\widetilde O(k/\varepsilon)$ 降低到 $\widetilde O(\sqrt{k}/\varepsilon)$。
英文摘要
We study differentially private learning problems in the realizable setting, where a hypothesis is specified by $k$ components. A direct iteration of private component learners is obstructed by a simple difficulty: an approximate choice of the next component may destroy exact realizability of the labeled sample, even when the next component is locally accurate. We restore realizability using the LabelBoost procedure of Beimel, Nissim, and Stemmer [SODA '15, Algorithmica '21] and recycle data through two alternating reservoirs. The resulting learner, for a target privacy $\varepsilon$, pays only $\widetilde O(\sqrt{k}/\varepsilon)$ overhead relative to the active sample requirement of a single component learning step at target accuracy $Θ(α/k)$. For learning $d$-dimensional halfspaces over a finite coordinate grid of size $L$, exact realizability makes the direct component-depth objective quasi-concave. Instantiating the framework with the IPConcave algorithm of Nissim, Tsfadia, and Yan [SODA '26] and with the quasi-concave optimizer of Cohen, Lyu, Nelson, Sarl'os, and Stemmer [STOC '23] yields a realizable sample complexity of $\widetilde{O}\left(\frac{1}{\varepsilon α}\cdot \min\{d^{2.5} \log^*L, \:\: d^{2.5} + d^{1.5} 2^{\log^*L}\}\right),$ which improves on the previously known bound of $\widetilde{O}\left(\frac{1}{\varepsilon α}\cdot\min\{\frac{1}α\cdot d^{5.5}\log^*L,\:\: d^{2.5}2^{\log^*L}\}\right).$ We also apply the framework to Boolean compositions: given proper private learners for classes $H_1,\ldots,H_k$, we obtain a proper private learner for $G(H_1,\ldots,H_k)$ for any fixed Boolean function $G:\{0,1\}^k\to\{0,1\}$. Compared with the closure theorem of Alon, Beimel, Moran, and Stemmer [COLT '20], this reduces the overhead on a common component sample bound from $\widetilde O(k/\varepsilon)$ to $\widetilde O(\sqrt{k}/\varepsilon)$.