arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06460cs.LG

掩码预训练(MPT)的理论框架

A Theoretical Framework for Masked Pretraining (MPT)

  • Peking University(北京大学)
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Qi Zhang, Runyu Zhou, Yifei Wang, Yisen Wang

AI总结:

本文提出掩码预训练的理论框架,证明掩码隐式创建正样本对,指出维度坍缩问题并提出U-MPT损失及新掩码策略,显著提升下游任务性能。

AI中文摘要:

近年来,基于重建预训练任务的掩码预训练(MPT)已成为跨多个领域的一种有前景的自监督学习范式,并在多种下游任务中取得了显著性能。然而,对MPT背后工作机制的理论理解仍然有限。在本文中,我们引入了一个新的理论框架来分析MPT,并理解掩码在提取有意义表示中的关键作用。我们在MPT与另一种流行的自监督范式——对比学习之间建立了理论联系。我们证明,掩码技术隐式地创建了语义相似的正样本对,并且重建损失在特征空间中将它们拉近。此外,作为隐式对齐的结果,我们指出了MPT的维度坍缩问题,并提出了一种均匀性增强的MPT(U-MPT)损失,该损失能有效解决此问题,并在下游任务中带来显著改进,包括线性评估、跨数据集微调以及在真实世界数据集上的分布外泛化。进一步地,我们建立了U-MPT的下游保证,并从理论上分析了掩码策略的影响。基于理论分析,我们提出了一种新的掩码策略,该策略提升了MPT的下游性能,并用我们的理论视角解释了当前掩码策略的改进。

英文摘要:

Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.

↑