arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19491cs.LGcs.CLmath.OCstat.ML

DeltaMomentum:基于键值结构的Delta规则各向异性动量更新

Activation-Keyed Momentum: An Anisotropic Momentum Update via the Delta Rule

发表机构卡内基梅隆大学
查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Euijin Hong, Guannan Qu

首次发表
浏览论文内容

中文总结 AI 辅助

DeltaMomentum将方向感知融入动量更新,可作为优化器动量缓冲器的即插即用替代,在预训练中减少AdamW步数,在多模型和数据集上验证了其有效性。

中文摘要 AI 辅助

大多数现代优化器将其动量构建为过去梯度的指数移动平均(EMA),以单一固定速率遗忘所有方向。然而,深度网络在训练过程中接收的输入可能具有高度各向异性,少数方向被频繁查询,而大多数方向很少被看到。近期方法通过在该缓冲器周围添加额外处理来解决这种各向异性,却未改变动量更新本身。我们提出DeltaMomentum,它将方向感知融入动量更新规则。核心观察是线性层的梯度可拆分为作为键的输入和作为值的输出侧误差。利用键值结构,DeltaMomentum通过标准Delta规则更新动量缓冲器,使每个方向的遗忘速率由其出现频率决定。我们证明它是有效动量,无需矩阵逆即可应用输入侧曲率校正,且在固定和漂移最优解下,比EMA更快清除陈旧方向。它可作为任何优化器动量缓冲器的即插即用替代,其系数在μP下可跨宽度迁移,额外计算量仅为门控MLP块线性成本的22.2%至25.0%,且无需持久内存。在FineWeb-Edu预训练中,采用DeltaMomentum的AdamW(DeltaAdamW)在67M参数规模下,最多可减少46.39±4.32%的步数达到AdamW的验证损失,在370M参数规模下减少22.12±0.80%的步数,且在Chinchilla最优预算的1B参数规模上,该优势依然存在。在相同协议下调优的Muon基线在两种语言模型规模上均优于DeltaAdamW,且该优势在SGD、ResNet-18和CIFAR-10上的ViT-Tiny模型中也成立。训练时的诊断结果证实了所预测机制:更好的梯度跟踪和更健康的输入方向。

英文摘要

Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic, with a few directions queried frequently while most are seen rarely. Preconditioning methods address this anisotropy by wrapping extra processing around this buffer and leave the momentum update itself unchanged. We propose Activation-Keyed Momentum (AK-Momentum), which builds direction-awareness into the momentum update rule. The gradient of a linear layer splits into an input activation that acts as a key and an output-side error that acts as a value. Keying on that activation, AK-Momentum updates the momentum buffer by the canonical delta rule, so each direction is forgotten at a rate set by how often it appears. We prove that it is a valid momentum, that it applies the input-side curvature correction without matrix inversion, and that it clears stale directions faster than EMA under both a fixed and a drifting optimum. It is a drop-in replacement for the momentum buffer of any optimizer, its coefficient transfers across widths under $μ$P, and its extra compute stays between $22.2\%$ and $25.0\%$ of a gated-MLP block's linear cost with no persistent memory. In FineWeb-Edu pretraining, AdamW with AK-Momentum (AK-AdamW) reaches AdamW's validation loss in up to $46.39 \pm 4.32\%$ fewer steps at 67M and $22.12 \pm 0.80\%$ at 370M over three seeds, and the gain persists at 1B on a Chinchilla-optimal budget. A Muon baseline tuned under the same protocol sits above AK-AdamW at both language-model scales, and the gain holds for SGD, ResNet-18, and ViT-Tiny on CIFAR-10. Training-time diagnostics confirm the predicted mechanism, better gradient tracking and healthier input directions.

补充信息

↑