坐标式Adam中的二阶矩记忆
Second-Moment Memory in Coordinatewise Adam
浏览论文内容
中文总结 AI 辅助
该研究针对坐标式Adam,证明其二阶矩记忆会在有限方差随机梯度下抑制优化收敛,给出了相关方向界与平稳性下界,指出长二阶矩记忆会减缓优化。
中文摘要 AI 辅助
Adam在其分母中保留过去平方梯度的移动平均,但这种记忆的优化成本尚未被充分理解。我们表明,即使在有限方差的随机梯度下,二阶矩记忆本身也会抑制向最优解的收敛。对于简单的两点神谕,在初始化瞬态后,预期的正归一化更新为$O(M_2^{-1/2})$,其中$M_2=(1-\beta_2)^{-1}$是二阶矩记忆长度。在所述记忆和步长缩放下,我们将此方向界转换为具有归一化间隙、平滑性和方差的光滑凸问题上同阶的平均平稳性下界。即使梯度噪声具有有限方差,长二阶矩记忆也会减缓优化过程。
英文摘要
Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood. We show that second-moment memory can itself suppress progress toward the optimum even under finite-variance stochastic gradients. For a simple two-point oracle, the expected positive normalized update is $O(M_2^{-1/2})$ after an initialization transient, where $M_2=(1-β_2)^{-1}$ is the second-moment memory length. We convert this directional bound, under the stated memory and stepsize scaling, into an average-stationarity lower bound of the same order on a smooth convex problem with normalized gap, smoothness, and variance. Long second-moment memory can slow optimization even when the gradient noise has finite variance.