大初始化下逼近逻辑斯蒂梯度下降轨迹的理想路径
Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization
浏览论文内容
中文总结 AI 辅助
本文提出一种几何近似方法,构造大初始化下逻辑斯蒂梯度下降轨迹的理想路径,证明其收敛性并给出损失渐近公式,揭示峰值评估损失可随初始化规模线性增长。
中文摘要 AI 辅助
现代训练在开始新任务时,通常从先前训练好的模型而非从零开始,这引发了初始化如何影响后续训练轨迹的问题。经典的隐式偏差结果刻画了长时间训练所选择的方向,但仅凭该方向无法提供关于中间行为的信息。我们通过几何近似全批量逻辑斯蒂梯度下降(GD)在严格线性可分数据上的轨迹来解决此问题,其中初始化规模为$R$,且受先前训练的启发。从任意极限归一化初始位置出发,我们使用最小范数投影规则构造一条由有限个线性段组成的唯一连续理想路径。该路径包含两个阶段:负间隔修正和最小间隔增长。我们证明,在显式的两阶段时间重参数化后,固定步长GD轨迹除以$R$在$R\to\infty$时于每个固定参数区间上一致收敛到该路径。此外,我们的定量误差界考虑了初始化扰动和阶段间的过渡。该近似提供了峰值评估损失和累积训练损失的渐近公式。特别地,即使两端损失均趋于零,峰值评估损失也可随$R$线性增长。修正和间隔增长阶段的累积损失分别按$R^2$和$R$归一化后,收敛到显式极限。在受控几何和固定图像特征上的实验补充了我们的理论结果。
英文摘要
Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale $R$ motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by $R$ converges uniformly to this path on every fixed parameter interval as $R\to\infty$. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in $R$ even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by $R^2$ and $R$, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.
发表机构
- Peking University(北京大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。