arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度学习优化中作为残差驱动乘子校正的动量

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

Zhixin Ren, Yao Lyu, Congrong Li, Liping Zhang, Shengbo Eben Li

arXiv 2608.12925首次发表:更新:

AI 中文总结

本研究提出AIM框架将动量解释为残差驱动乘子校正,并基于此开发RADAR优化器,经多任务实验验证其性能优于现有强自适应优化器基线。

AI 中文摘要

基于动量的优化器在现代深度学习中被广泛应用,但动量递推、更新几何与加速之间的关系仍仅被部分理解。我们基于残差惩罚变量拆分开发了ADMM启发式动量(AIM)框架,该框架将动量解释为由拆分残差驱动的类乘子校正。AIM从ADMM式乘子更新中恢复梯度的指数移动平均,并分离出实际优化器中通常交织的两种机制:残差惩罚决定更新几何,而与目标相关的子问题近似决定加速形式。基于AIM,我们提出带加速残差的相对论自适应梯度下降(RADAR),其结合了相对论自适应几何、解耦残差校正和二阶动量滤波以改进更新方向与动量估计。我们通过方差扰动李雅普诺夫漂移分析建立了随机收敛性。在监督视觉学习、语言建模和强化学习上的实验表明,RADAR相较于强自适应优化器基线取得了一致的性能提升。

英文摘要

Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain only partially understood. We develop an $\textbf{A}$DMM-$\textbf{I}$nspired $\textbf{M}$omentum (AIM) framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual. AIM recovers the exponential moving average of gradients from an ADMM-style multiplier update and separates two mechanisms that are usually intertwined in practical optimizers: the residual penalty determines the update geometry, whereas the approximation of the objective-related subproblem determines the acceleration form. Building on AIM, we propose $\textbf{R}$elativistic $\textbf{A}$daptive gradient $\textbf{D}$escent with $\textbf{A}$ccelerated $\textbf{R}$esidual (RADAR), which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and momentum estimation. We establish stochastic convergence through a variance-perturbed Lyapunov drift analysis. Experiments on supervised vision learning, language modeling, and reinforcement learning show that RADAR achieves consistent improvements over strong adaptive optimizer baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑