arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.05136cs.LG

解耦正交动力学:深度网络优化器的正则化

Decoupled Orthogonal Dynamics: Regularization for Deep Network Optimizers

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao Chen, Jinghui Yuan, Hanmin Zhang

更新

AI总结:

AdamO通过解耦幅度和方向的动力学,改进深度学习优化器的泛化和稳定性。

AI中文摘要:

标准的AdamW中的权重衰减是否真正最优?尽管AdamW将权重衰减与自适应梯度缩放解耦,但仍然存在根本性的冲突:径向拔河。在深度学习中,梯度倾向于增加参数范数以扩展有效容量,同时引导方向以学习特征,而权重衰减无差别地抑制范数增长。这种推拉相互作用导致径向振荡,向Adam的二阶矩估计注入噪声,可能破坏精细的切向特征学习。我们主张幅度和方向扮演不同的角色,并应在优化器动力学中解耦。我们提出了正交动力学解耦,并将其实例化为AdamO:一种SGD风格的更新处理一维范数控制,而Adam的自适应预条件被限制在切向子空间中。AdamO进一步结合了曲率自适应的径向步长调整和架构感知的规则和投影,用于尺度不变层和低维参数。在视觉和语言任务上的实验表明,AdamO在不引入额外复杂约束的情况下,比AdamW在泛化和稳定性方面有所改进。

英文摘要:

Is the standard weight decay in AdamW truly optimal? Although AdamW decouples weight decay from adaptive gradient scaling, a fundamental conflict remains: the Radial Tug-of-War. In deep learning, gradients tend to increase parameter norms to expand effective capacity while steering directions to learn features, whereas weight decay indiscriminately suppresses norm growth. This push--pull interaction induces radial oscillations, injecting noise into Adam's second-moment estimates and potentially degrading delicate tangential feature learning. We argue that magnitude and direction play distinct roles and should be decoupled in optimizer dynamics. We propose Orthogonal Dynamics Decoupling and instantiate it as AdamO: an SGD-style update handles the one-dimensional norm control, while Adam's adaptive preconditioning is confined to the tangential subspace. AdamO further incorporates curvature-adaptive radial step sizing and architecture-aware rules and projections for scale-invariant layers and low-dimensional parameters. Experiments on vision and language tasks show that AdamO improves generalization and stability over AdamW without introducing additional complex constraints.

↑