发表机构
Technical University of Munich; Columbia University(慕尼黑工业大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析了自监督预训练中相关样本的增强汇集问题,发现其误差界不劣于基线划分方法,部分场景下收敛更快,解释了大量增强的实践优势。
AI 中文摘要
自监督学习依赖于无标签数据点 $x$ 的所谓数据增强 $\u03d5(x)$——例如,对图像 $x$ 中的随机像素进行掩码——这些增强应保持 $x$ 的标签不变,且通常用于学习下游任务的低复杂度不变子空间 $\u2130$。在实践中,尽管同一数据点 $x$ 的不同增强 $\u03d5_l(x)、\u03d5_k(x)$ 之间存在明显的相互依赖关系,人们仍会将这些增强 $\u007b \u03d5_l(x_i) \u007d$ 汇集在一起来学习 $\u2130$。然而,该主题的理论研究通常考虑避免此类依赖的方法,因此仅限于在较小的独立数据子集上操作。本研究表明,尽管存在相互依赖,将增强汇集在一起仍是比将数据划分为独立数据子集的基线方法更好的选择。更准确地说,在估计 $\u2130$ 的背景下,汇集的统计估计误差界从未比基线划分方法差,且在某些情况下——例如在浅层神经网络上基于掩码或噪声注入的增强——简单汇集会在增强数量方面带来更快的收敛速率。当不同增强 $\u03d5_l(x)、\u03d5_k(x)$ 之间的相关性对估计的影响较小或有助于降低估计方差时,汇集的优势尤为显著。因此,该分析为自监督预训练中汇集增强样本的成功提供了新的见解,并为实践中倾向于使用大量增强的做法提供了直观解释。
英文摘要
Self-supervised learning relies on so-called data augmentations $ϕ(x)$ of unlabeled datapoints $x$ --- for example, masking random pixels in an image $x$ --- that should leave the label of $x$ invariant and are often used to learn a lower-complexity invariant subspace $\cal V$ for downstream tasks. In practice, such augmentations $\{ ϕ_l(x_i) \}$ are pooled together to learn $\cal V$, despite obvious inter-dependencies between different augmentations $ϕ_l(x), ϕ_k(x)$ of the same datapoint $x$. However, theoretical works on the subject typically consider procedures that avoid such dependencies, and are therefore limited to operate on smaller subsets of independent data. We show in this work that pooling augmentations together, despite inter-dependencies, is a better alternative than the baseline of partitioning the data into subsets of independent data. More precisely, in the context of estimating $\cal V$, the statistical estimation error bounds for pooling are never worse than the partitioning baseline, and in some cases --- such as masking or noise injection-based augmentations over a shallow neural network --- naive pooling leads to faster rates in terms of the number of augmentations. The benefits of pooling are particularly prominent when the correlations between different augmentations $ϕ_l(x), ϕ_k(x)$ have mild effects on estimation or help decrease the estimation variance. The analysis, therefore, yields new insights into the success of pooling augmented samples in self-supervised pre-training, and provides an intuition behind the practical preference towards using many augmentations.