arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为什么自适应优化器低估稀有词元

Why Adaptive Optimizers Underestimate Rare Tokens

Sangsidhya Kar

arXiv 2609.37535首次发表:更新:

发表机构

Presidency University, Kolkata(加尔各答总统大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文分析自适应优化器(如Adam、RMSProp)对稀有词元产生低估偏差的机制,指出其归一化更新导致训练不动点偏移,并推导出偏差的闭式表达式,实验验证了该现象。

AI 中文摘要

在softmax输出层中,稀有词元在大多数步骤中接收到较小的正logit梯度,而在作为目标的少数步骤中接收到大得多的负梯度。SGD只是将这些贡献相加。而逐坐标的自适应方法(如Adam、RMSProp和符号下降)则会将每次更新除以运行中的幅度估计值,且该估计值在词元出现后立即达到最大。这种不平衡产生两种影响。在整个输出层层面,我们刻画了哪些优化器能保持平均输出嵌入:所有更新为过去梯度的线性函数的方法,以及Kronecker分解和正交化方法(如Shampoo和Muon)都能保持。Adam、Adafactor、Lion和符号下降则不能,对于这些方法,我们获得了变化的逐步精确表达式。在单个稀有词元层面,相同的归一化会移动训练不动点。在unigram模型中,符号下降以恒定的期望速率降低每个出现频率低于半数小批次的词元的logit。对于周期性到达的RMSProp,我们可以闭式求解不动点:如果词元至少连续两个小批次未出现,则其平衡概率严格低于其数据频率(对任何学习率均如此),且比值趋向于κ/(2(e^{κ/2}-1))。这里κ是平均出现间隔步数除以二阶矩时间常数1/(1-β2)。在同一模型中,SGD和AMSGrad保持无偏不动点。我们在unigram模型和从已知生成分布训练的小型语言模型中测试了这些预测。在随机到达情况下,偏差大于周期性公式的预测;在语言模型中,具有有偏不动点的优化器对生成分布的拟合也较差。

英文摘要

In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a running estimate of its magnitude, and that estimate is largest immediately after the token appears. This imbalance has two effects. At the level of the whole output layer, we characterize which optimizers preserve the mean output embedding: every method whose update is linear in past gradients does, as do Kronecker-factored and orthogonalized methods such as Shampoo and Muon. Adam, Adafactor, Lion, and sign descent do not, and for these methods we obtain an exact step-by-step expression for the change. At the level of an individual rare token, the same normalization shifts the training fixed point. In the unigram model, sign descent lowers the logit of every token that occurs in fewer than half of the minibatches at a constant expected rate. For RMSProp with periodic arrivals, we can solve the fixed point in closed form: if a token is absent for at least two consecutive minibatches, its equilibrium probability is strictly below its data frequency for every learning rate, and the ratio tends to $κ/(2(e^{κ/2}-1))$. Here $κ$ is the mean number of steps between occurrences divided by the second-moment time constant $1/(1-β_2)$. In the same model, SGD and AMSGrad retain the unbiased fixed point. We test these predictions both in a unigram model and in a small language model trained from a known generating distribution. With random arrivals, the bias is larger than the periodic formula predicts; in the language model, the optimizers with the biased fixed point also fit the generating distribution less well.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑