发表机构
University of Alberta(阿尔伯塔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文定量刻画了Adam在可分离线性分类中从最大范数间隔到欧几里得间隔的转变,揭示了步长衰减指数对转变时间和泛化影响的精确规律。
AI 中文摘要
我们考虑确定性、全批次、偏差校正的Adam在softmax参数化下的可分离线性分类中的行为,损失函数为对数损失。在此设置下,在广泛条件下,已知当稳定性常数$\epsilon$为零时,Adam趋向于最大范数间隔最优性,而当$\epsilon$为正时,已知其趋向于欧几里得间隔最优性。我们的主要贡献是对小固定正$\epsilon$下Adam行为的定量描述。我们给出了充分条件,使得Adam训练的分类器在更新变得类似梯度之前几乎最大化最大范数间隔。我们还表明,分类器仅在更晚的时候达到固定的目标欧几里得间隔。具体来说,我们证明对于指数为$a$的多项式递减步长,其中$1/3<a<1$,更新在$\Theta(\log(1/\epsilon)^{1/(1-a)})$次迭代后近似与负梯度成比例。此时,分类器仍几乎最大化最大范数间隔。达到一个高于所有最大范数最优分类器但低于最优的固定目标欧几里得间隔,被证明需要$\epsilon^{-\Theta(1)/(1-a)}$次迭代。在逆线性步长衰减($a=1$)下,更新转变需要多项式多次迭代,而达到目标间隔则需要指数多次迭代。实验支持这些预测。分类器后来的变化可以在训练误差达到零后改善或恶化泛化,将分析与grokking及其逆过程联系起来。
英文摘要
We consider the behavior of deterministic, full-batch, bias-corrected Adam in separable linear classification with softmax parametrization under log-loss. In this setting, under a wide range of conditions Adam is known to approach max-norm-margin optimality when its stability constant $ε$ is zero, while with a positive $ε$, it is known to approach Euclidean-margin optimality. Our main contribution is the quantitative description of Adam's behavior for small fixed positive $ε$. We give sufficient conditions under which an Adam-trained classifier nearly maximizes the max-norm margin before the updates become gradient-like. We also show that the classifier reaches a fixed target Euclidean margin only much later. Specifically, we show that for polynomially decreasing stepsizes with exponent \(a\), where \(1/3<a<1\), the updates become approximately proportional to the negative gradient after $Θ(\log(1/ε)^{1/(1-a)})$ iterations. At that time, the classifier still nearly maximizes the max-norm margin. Reaching a fixed target Euclidean margin above that of every max-norm-optimal classifier, but below the optimum, is shown to require $ε^{-Θ(1)/(1-a)}$ iterations. Under inverse-linear stepsize decay (\(a=1\)), the update transition takes polynomially many iterations, whereas reaching the target margin takes exponentially many. Experiments support these predictions. The later change in the classifier can improve or worsen generalization after training error reaches zero, connecting the analysis to grokking and its reverse.