arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21815cs.LGcs.AI

Matrix AdaGrad:行式与列式自适应次梯度方法

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenpeng Zhang, Runsheng Yu, Peilin Zhao

中文总结 AI 辅助

本文提出行式与列式矩阵AdaGrad优化方法,通过在线镜像下降框架利用矩阵结构,实现更紧的遗憾界,并在矩阵分解和深度网络训练中提升稳定性与可训练性。

中文摘要 AI 辅助

自适应优化方法如AdaGrad和Adam在现代神经网络训练中被广泛使用,但其自适应缩放主要针对向量值参数设计,并未显式利用矩阵结构。最近的矩阵感知优化器展示了结构化优化的优势,然而,缺乏一个通用的理论框架来推导与AdaGrad相当的矩阵感知自适应性。在本工作中,我们开发了一个通用的在线镜像下降框架,针对矩阵值参数采用自适应近端函数,为通过在线遗憾最小化推导矩阵感知自适应优化提供了原则性方法。通过引入行式和列式矩阵近端函数并分析由此产生的遗憾权衡,我们推导出行式矩阵AdaGrad(Row-AdaGrad)和列式矩阵AdaGrad(Column-AdaGrad),其自适应缩放由累积的行式或列式梯度范数决定。我们建立了遗憾保证,并表明在结构化梯度下,这些矩阵感知的界可以严格地比逐元素AdaGrad的界更紧。在矩阵分解和深度神经网络训练上的实验进一步证明了将自适应缩放与矩阵结构对齐的益处,包括在更大学习率和更深网络深度下改善优化稳定性和可训练性。

英文摘要

Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop an Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled methodology for deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions, our framework explicitly reveals the trade-off governing adaptive scaling: increasing the scaling factors reduces the gradient-dependent dual norm term while increasing the cost of evolving the proximal geometry. In the row-wise setting, this trade-off becomes separable under diagonal parameterization, allowing the adaptive scaling for each row to be derived independently by minimizing its corresponding row-wise regret bound. The column-wise counterpart follows directly by applying the row-wise construction to the transposed matrix. This framework yields Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad as concrete instantiations, with regret guarantees that are strictly tighter than those of entry-wise AdaGrad under row-sparse or column-sparse gradient structures. Experiments on matrix factorization and stacked deep MLP training further demonstrate the benefits of matrix-aware adaptive scaling, yielding improved optimization performance in both settings and enhanced optimization stability and trainability at larger learning rates and greater network depths in the latter.

↑