发表机构
School of the Gifted Young, University of Science and Technology of China; Institute of Robotics and Automatic Information Systems, College of Artificial Intelligence, Nankai University; National Key Lab of General AI, School of Intelligence Science and Technology, Peking University(中国科学技术大学少年班; 南开大学人工智能学院机器人与自动信息系统研究所; 北京大学智能科学与技术学院通用人工智能全国重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出统一理论框架,首次证明SOAP和Adam的收敛速率,并改善维度依赖性,涵盖多种基于旋转的矩阵优化器。
AI 中文摘要
在这项工作中,我们开发了一个统一的理论框架,用于分析基于旋转的矩阵优化器的收敛性。这类优化器应用正交变换将动量映射到旋转空间中,在该空间中进行逐坐标或归一化更新,然后再将更新旋转回原始空间。我们的框架涵盖了著名的矩阵优化器,包括SOAP、Conda和截断的SPlus,以及作为特殊情况的矩阵参数化Adam。作为主要结果,我们首次建立了SOAP的收敛速率,具有尖锐的维度依赖性,以及以核范数度量的Adam的收敛速率。我们的框架还涵盖了基于旋转的优化器的一个更一般的变体,该变体允许任意正交旋转矩阵,并允许这些矩阵在每次迭代时依赖于当前的随机梯度。在技术上,我们的框架利用了一种行列迹控制论证,将逐元素界转化为两个对角控制矩阵上的界,从而改善了所得界的维度依赖性。
英文摘要
In this work, we develop a unified theoretical framework for analyzing the convergence of rotation-based matrix optimizers, which apply orthogonal transformations to map the momentum into a rotated space, perform coordinate-wise or normalized updates there, and then rotate the updates back. Our framework encompasses prominent matrix optimizers, including SOAP, Conda, and truncated SPlus, as well as matrix-parameterized Adam as a special case. As our main result, we establish, for the first time, the convergence rate of SOAP with sharp dimensional dependence, as well as the convergence rate of Adam measured by the nuclear norm. Our framework also covers a more general variant of rotation-based optimizers that allows arbitrary orthogonal rotation matrices, allowing these matrices to depend on the current stochastic gradient at each iteration. Technically, our framework leverages a row-column trace-control argument that converts elementwise bounds into bounds on two diagonal control matrices, thereby improving the dimension dependence of the resulting bound.
CommentsThis paper subsumes our prior manuscript: Convergence Rate Analysis of SOAP with Arbitrary Orthogonal Projection Matrices, arXiv:2604.21616