发表机构
McGill University; Mila, Quebec AI Institute(麦吉尔大学; 魁北克人工智能研究所米拉开创)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自监督学习标记样本效率缺乏理论解释的问题,通过将留一法稳定性工具用于数据增强图,证明快速转导速率\(O(1/n_L)\),明确增强质量,解释准确率与标签数量曲线,且使用简化损失恢复理想特征。
AI 中文摘要
自监督学习能以少量标签匹配监督学习的准确率,但其背后的标记样本效率缺乏理论解释,本文给出了解释。数据增强在未标记数据上诱导出相似性图,下游在该图上的学习是图拉普拉斯正则化学习。通过将约翰逊和张的留一法稳定性工具应用于增强图,证明了快速转导速率为\(O(1/n_L)\)(标签数量),取代了监督学习的\(O(1/\sqrt{n_L})\),且无需基于极限分析的不切实际假设。该界限明确了增强质量:预期误差至多为\(C/n_L + R_{\mathrm{DA}}(y)\),其中数据增强对齐误差\(R_{\mathrm{DA}}(y)\)是跨越标签边界的增强的图割质量,所以良好的增强使少量标签就足够。分析使用了简化损失,在无限数据极限下仍能恢复前\(K\)个理想特征,即翟等人研究的增强核特征空间。结果解释了观察到的准确率与标签数量曲线,而非仅界定泛化差距。
英文摘要
Self-supervised learning matches supervised accuracy from a fraction of the labels, but the labeled-sample efficiency behind this has lacked a theoretical explanation. We provide one. Data augmentation induces a similarity graph on the unlabeled data, so downstream learning on that graph is graph-Laplacian-regularized learning. We prove a fast transductive rate, $O(1/n_L)$ in the number of labels, in place of the supervised $O(1/\sqrt{n_L})$, by carrying the leave-one-out stability apparatus of Johnson and Zhang (JMLR 2007) over to the augmentation graph, and without the unrealistic assumptions of limit-based analyses (exact kernel, generalizing features). The bound makes augmentation quality explicit: the expected error is at most $C/n_L + R_{\mathrm{DA}}(y)$, where the data-augmentation alignment error $R_{\mathrm{DA}}(y)$ is proportional to the graph-cut mass of augmentations that cross a label boundary, so good augmentations let few labels suffice. The analysis uses a streamlined loss that drops the projector, negative-sample, and orthogonality overhead of standard objectives yet still recovers the top-$K$ ideal features in the infinite-data limit, the augmentation-kernel eigenspace studied by Zhai et al. The bound gives a mechanistic account of the accuracy-versus-label-count curve through augmentation quality, verified in a controlled model where the constants are known.
Comments23 pages, 2 figures