发表机构
Memorial University of Newfoundland(纽芬兰纪念大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示梯度下降隐式偏向低秩解而Adam不偏向的原因是损失的规范对称性,确定了等变优化器的条件,通过实验分析了9种优化器的性能,指出基选择是优化器的关键决策而非调优细节。
AI 中文摘要
在因子化模型 $W = UV^\top$ 上的梯度下降会隐式偏向低秩解,而从相同小初始化开始的Adam则不会。我们将这种差异归因于损失的规范对称性,即损失在 $(U, V) \mapsto (UQ, VQ)$ 变换下保持不变。梯度流的低秩机制仅对规范等变的优化器可用,该条件是迁移的必要条件但非低秩恢复的充分条件。梯度下降、带动量的优化器、“共享标量”Adam、Muon和Shampoo满足该条件,而Adam、RMSProp及其他逐坐标方法不满足。一个结构定理将无记忆等变规则刻画为恰好是格拉姆矩阵确定的左预条件子,一个迁移定理将梯度流的路径性质传递给共享标量流。随后,我们在欠定矩阵感知任务上对9种更新规则,按相对于植入真值的恢复误差进行排序。从逐坐标到共享标量预条件的单参数族会单调恢复该偏向,从而确定各向异性是原因。“谱调度”调和了关于Muon的两份矛盾报告:等速率更新可精确恢复低秩目标,但随谱尾部增长会失去优势。在Transformer中,Adam在第一步就分离了两个规范等价的初始化,而等变优化器则保持在浮点精度,最终得到的头级不变量 $W_Q^\top W_K$ 的相对弗罗贝尼乌斯距离相差56%,该差距无法通过任何头级旋转消除。在两个高光谱数据集上,当训练损失匹配时,梯度下降在最低采样密度下将保留误差降低了43-44%,且有效秩更低。因此,基选择并非调优细节,而是优化器选择哪种插值函数的决策。
英文摘要
Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mechanism is available to an optimizer only if that optimizer is gauge-equivariant, a condition necessary for the transfer but not sufficient for low-rank recovery. Gradient descent, momentum, "shared-scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate-wise methods do not. A structure theorem characterizes the memoryless equivariant rules as exactly the Gram-determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common-scalar flows. We then sort nine update rules on underdetermined matrix sensing by recovery error against the planted ground truth. A one-parameter family from coordinate-wise to shared-scalar preconditioning restores the bias monotonically, isolating anisotropy as the cause. A "spectral schedule" reconciles two opposing reports about Muon: equal-rate updates recover exactly low-rank targets but lose their edge as the spectral tail grows. In transformers, Adam separates two gauge-equivalent initializations at the first step, where the equivariant optimizers stay at float precision, and ends with the per-head invariants $W_Q^\top W_K$ 56% apart in relative Frobenius distance, a gap no per-head rotation can close. On two hyperspectral datasets at matched training loss, gradient descent cuts held-out error by 43-44% at the lowest sampling density, and at lower effective rank. Basis choice is therefore not a tuning detail but a decision about which interpolant the optimizer selects.
Comments22 pages main text + appendices, 5 figures. Code, seeds, and raw run records: https://github.com/idevender/loss-basis-adam