发表机构
Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过统一框架研究优化器几何与深度如何共同控制矩阵分解中的谱学习动力学,推导多种优化器的模态演化定律,揭示归一化、累积状态、阻尼和深度对模态出现、持续和消失的调控作用。
AI 中文摘要
基于矩阵和曲率的优化器近期取得的成功,重新激发了人们对更新几何如何塑造学习的兴趣。这些方法对更新进行归一化或预处理,从而改变训练过程中不同分量的进展方式。在深度矩阵分解中,减缓梯度下降(GD)的几何通过延迟小奇异模式的涌现而有利于低秩解。这引出一个问题:当归一化或曲率校正削弱或消除这种减缓时,这种谱偏差还剩下什么。我们通过一个统一的奇异模式动力学框架,研究优化器几何和深度如何共同控制这一行为。在显式平衡和对齐假设下,我们推导了欧几里得GD、SignGD的逐坐标更新和瞬时Adam近似、Muon和累积Shampoo的谱更新,以及K-FAC的块曲率模型的出生、饱和和衰减定律。由此产生的图景并非简单的从强到弱的低秩偏差排序。需要强调的是,SignGD和理想Muon消除了发散的出生障碍,并在有限时间内将不支持的模态驱动到零。累积Shampoo最初保留了GD的深度相关障碍,随后累积梯度产生追赶阶段,同时使先前活跃的模态越来越持久。无阻尼K-FAC抵消了因分解引起的减缓,同时保留了目标奇异值的排序,而正阻尼引入了谱阈值,低于该阈值,慢GD相位定律重新出现。这些结果为归一化、累积状态、阻尼和深度提供了直接解释,作为决定模态在训练过程中何时出现、持续和消失的控制手段。
英文摘要
Recent successes of matrix- and curvature-based optimizers have renewed interest in how update geometry shapes learning. These methods normalize or precondition updates, changing how different components progress during training. In deep matrix factorization, the geometry that slows gradient descent (GD) favors low-rank solutions by delaying the emergence of small singular modes. This raises the question of what remains of that spectral bias when normalization or curvature correction weakens or removes the slowdown. We study how optimizer geometry and depth jointly govern this behavior through a common framework for singular mode dynamics. Under explicit balance and alignment assumptions, we derive birth, saturation, and decay laws for Euclidean GD, coordinate-wise updates of SignGD and an instantaneous Adam approximation, spectral updates of Muon and cumulative Shampoo, and a block curvature model of K-FAC. The resulting picture is not a simple ordering from stronger to weaker low-rank bias. To highlight, SignGD and ideal Muon eliminate the divergent birth barrier and drive unsupported modes to zero in finite time. Cumulative Shampoo initially retains GD's depth-dependent barrier, then accumulated gradients produce a catch-up phase while making previously active modes increasingly persistent. Undamped K-FAC cancels the factorization-induced slowdown while preserving the ordering of the target singular values, whereas positive damping introduces a spectral threshold below which the slow GD phase laws reappear. These results give normalization, accumulated state, damping, and depth a direct interpretation as controls determining when modes emerge, persist, and disappear during training.
CommentsUnder review