arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34667stat.MLcs.LG

双时间尺度微调可证明学习两层ReLU网络的新特征

Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

Etienne Boursier, Nicolas Flammarion

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出双时间尺度微调方法,证明其能从稀缺数据中学习新特征,同时保留预训练特征,且所需样本量独立于预训练特征数,优于随机初始化。

中文摘要 AI 辅助

在专门任务上使用稀缺数据微调预训练模型是现代深度学习的核心。尽管其经验成功,但微调的理论理解仍然有限。我们引入了一个高斯多指标设置来研究从预训练权重开始的微调,其中教师网络具有$m+1$个特征,其中$m$个在预训练期间学习,一个必须在微调期间学习。对于两层ReLU网络,我们证明双时间尺度训练,即外部权重的更新速度无限快于隐藏权重,能够学习新的任务特定特征,同时保留模型表示中的预训练特征。此外,这种恢复仅需要$\mathcal{O}(d)$个微调样本,与预训练特征的数量无关。相比之下,使用随机初始化,相同数量的样本不足以恢复目标参数。因此,我们的结果表明,预训练可以诱导一种隐式偏差,相对于随机初始化具有明显的统计优势,从而能够从稀缺的微调数据中学习特征。

英文摘要

Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.

发表机构

  • INRIA(法国国家信息与自动化研究所)
  • Université Paris-Saclay(巴黎萨克雷大学)
  • EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑