在脂肪破碎维度上近线性大小的平方损失不可知样本压缩方案
An Agnostic Sample Compression Scheme for Squared Loss of Near-Linear Size in the Fat-Shattering Dimension
浏览论文内容
中文总结 AI 辅助
针对平方损失,构造了大小与脂肪破碎维度近线性、与样本量无关的不可知压缩方案,移除对偶因子,解决了开放问题。
中文摘要 AI 辅助
我们针对每个函数类 $\mathcal{F}\subseteq[0,1]^{\mathcal{X}}$ 和每个精度 $0<\alpha\le 1$,构造了一个用于经验平方损失的不可知样本压缩方案:对于任意带有任意(含噪声)标签的有限样本 $S\in(\mathcal{X}\times[0,1])^m$,该方案仅存储至多 $O(\mathrm{fat}(\mathcal{F},c'\alpha)\cdot\log^3(2/\alpha))$ 个原始带标签样本和辅助比特,与样本大小 $m$ 无关,并能重构一个函数 $\hat f$,使得 $L_2(\hat f,S)\le\inf_{f\in\mathcal{F}}L_2(f,S)+\alpha$。这肯定地解决了 Attias、Hanneke、Kontorovich 和 Sadigurschi(ICML 2024,第5节)提出的开放问题,该问题要求一个大小为 $\mathrm{fat}(\mathcal{F},c\alpha)\cdot\mathrm{polylog}(c/\alpha)$ 的不可知 $\ell_2$ 压缩方案。所有先前已知的有界大小构造,无论是不可知的还是甚至可实现的,都会引入一个乘性的对偶脂肪破碎因子,该因子可能比原始维度指数级更大;我们的方案完全移除了对偶因子,包括在可实现的情况下。先前工作中的对偶因子仅通过一个强制在样本上均匀逼近的稀疏化步骤进入。通过仅针对样本点的 $(1-\epsilon)$ 比例(这对于有界范围内的平均损失保证已足够),K'egl 的提升边际界产生了 $O(\log(1/\epsilon))$ 轮,与 $m$ 无关,且永远不需要稀疏化。提升器的合成目标标签(接近最优的 $f^*\in\mathcal{F}$ 的值)通过附加在存储的原始示例上的量化侧信息比特进行传输,而平方损失的交叉项迫使弱学习尺度为 $\Theta(\alpha)$,与开放问题的同尺度形式相匹配。
英文摘要
We construct, for every function class $\mathcal{F}\subseteq[0,1]^{\mathcal{X}}$ and every accuracy $0<α\le 1$, an agnostic sample compression scheme for the empirical squared loss: for every finite sample $S\in(\mathcal{X}\times[0,1])^m$ with arbitrary (noisy) labels, the scheme stores at most $O(\mathrm{fat}(\mathcal{F},c'α)\cdot\log^3(2/α))$ original labeled examples and auxiliary bits, independent of the sample size $m$, and reconstructs a function $\hat f$ with $L_2(\hat f,S)\le\inf_{f\in\mathcal{F}}L_2(f,S)+α$. This resolves, in the positive, the open problem of Attias, Hanneke, Kontorovich, and Sadigurschi (ICML 2024, Section 5), which asks for an agnostic $\ell_2$ compression scheme of size $\mathrm{fat}(\mathcal{F},cα)\cdot\mathrm{polylog}(c/α)$. All previously known bounded-size constructions, agnostic and even realizable, incur a multiplicative dual fat-shattering factor, which can be exponentially larger than the primal dimension; our scheme removes the dual factor entirely, including in the realizable case. The dual factor in prior work enters solely through a sparsification step that forces uniform approximation on the sample. By targeting only a $(1-ε)$-fraction of sample points, which suffices for an average-loss guarantee over a bounded range, K'egl's boosting margin bound yields $O(\log(1/ε))$ rounds independent of $m$, and sparsification is never needed. The booster's synthetic target labels (values of a near-optimal $f^*\in\mathcal{F}$) are transmitted through quantized side-information bits attached to stored original examples, and the cross term of the squared loss forces the weak-learning scale $Θ(α)$, matching the same-scale form of the open problem.