arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

掩码重采样的隐藏优势:掩码自编码器理论

The hidden advantage of mask resampling: a theory of masked autoencoders

Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová

arXiv 2610.01578首次发表:更新:

发表机构

École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过高维模型证明掩码线性重建在异质噪声下能以线性样本复杂度恢复潜在特征,优于PCA,并揭示掩码重采样多样性可降低样本复杂度,实验表明动态掩码在CNN和视觉变换器中具下游优势。

AI 中文摘要

为什么掩码预测能够学习到未掩码重建所遗漏的有用表示?我们在一个高维掩码自编码器(MAE)模型中研究这一问题,该模型在具有共享潜在结构和异质噪声的数据上训练。我们证明,在未掩码线性重建(等价于主成分分析,PCA)失败的机制下,掩码线性重建能够以线性样本复杂度恢复潜在特征。该分析还量化了掩码重采样(掩码预训练的一个既定组成部分)的统计优势。通过为每个样本引入固定集合的$K$个掩码,我们刻画了其对特征恢复和下游性能的影响,识别出更大的掩码多样性降低样本复杂度的机制。在此预测的指导下,我们发现标准图像训练流程中的随机裁剪和翻转可能通过更新预测任务而掩盖掩码重采样的优势,即使补丁掩码固定。移除这些变换揭示了在CNN自编码器和视觉变换器中,动态掩码相比静态掩码具有下游优势。一项补充的BERT试点研究发现,在掩码多样性更大的情况下,下游语言任务受益。我们的结果将掩码预测目标的好处与掩码多样性的好处分开,并展示了可处理的理論如何指导实验,揭示标准训练实践所隐藏的优势。

英文摘要

Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of $K$ masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑