发表机构
EPFL; INRIA; Université Paris-Saclay(洛桑联邦理工学院; 法国国家信息与自动化研究所; 巴黎萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过稀疏线性回归和对角线性网络,证明从预训练初始化微调能利用预训练支持信息降低样本复杂度,其效果可与显式利用该信息的加权Lasso估计器相媲美。
AI 中文摘要
将预训练模型适应到数据有限的下游任务已成为现代深度学习中的核心范式。然而,尽管其在实践中取得了广泛成功,微调如何利用预训练信息在理论上仍鲜有研究。我们通过稀疏线性回归和两层对角线性网络的视角研究从预训练权重进行微调。在我们的设置中,预训练通过初始化预测器的支持(及符号)提供信息,该支持可能包含与下游任务相关的坐标。我们展示了预训练信息如何重塑隐式偏差和训练动态,从而能够降低恢复目标参数和支持的样本复杂度。特别地,对于具有正确继承符号的干净初始化,我们表明所需的样本量与加权Lasso估计器相当,后者通过适当选择的正则化器显式利用预训练支持。我们的结果因此展示了预训练权重中编码的信息如何被基于梯度的微调隐式利用,从而减少恢复下游任务所需的数据量。
英文摘要
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.