发表机构
University of Oxford; Shanghai Jiao Tong University(牛津大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在岭回归梯度下降框架下,将预训练-微调计算分配建模为两阶段优化问题,并基于谱分解刻画了数据相关评估几何下的最优计算分割。
AI 中文摘要
预训练后接着微调引入了一个计算分配问题:在固定的训练预算下,用于改进上游目标的计算会减少可用于下游适配的计算。尽管这一问题具有实际重要性,但即使在简单模型中,其理论理解仍不充分。本文将这种分配问题建模为在固定总优化预算下、采用两阶段预训练-微调过程时的计算分割问题,并以梯度下降训练的正则化最小二乘作为易处理的设定。我们刻画了由微调问题引起的数据相关评估几何下的最优分割。我们的结果表明,分配取决于预训练方向如何影响微调预测,以及微调变化如何通过下游数据几何被观察。特别地,相关量由预训练和微调经验协方差中与预测相关的谱分量决定。在技术上,分析依赖于基不变、特征空间级别的谱分解,以及对非交换的预训练和微调动力学的微扰控制。
英文摘要
Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast this allocation as a compute-split problem under a two-stage pretrain--fine-tune procedure with fixed total optimisation budget, using regularised least squares trained by gradient descent as a tractable setting. We characterise the optimal split under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.