arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从几何到泛化:为何行归一化能胜过Adam与Muon

From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon

Jihwan Kim, Dogyoon Song, Chulhee Yun

arXiv 2610.11309首次发表:更新:

发表机构

Seoul National University; KAIST; University of California, Davis(首尔国立大学; 韩国科学技术院; 加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究证明行归一化在高维多分类任务中泛化性能优于Adam和Muon,其逐类几何结构可保留决策边界方向,合成与语言模型实验验证了该优势。

AI 中文摘要

不同优化器在拟合相同训练数据时,会选择具有显著不同几何结构的分类器,但这种差异是否会对泛化性能产生可证明的影响仍不明确。我们表明,在高维多分类任务中,行归一化(row-wise normalization)能实现比全批Adam(full-batch Adam,作为随机重排Adam的替代)和精确SVD Muon更高的总体准确率。在各向同性高斯云数据模型下,这一优势源于行归一化的逐类欧氏几何渐近保留了总体决策边界方向,而Adam的坐标式几何与Muon的谱几何则会引入非零失真。超出各向同性假设后,该优势在针对类别均值的全批训练(类别均值与测试噪声协方差相互独立)中依然存在;对于类别均值指数小于1的幂律谱,即使在高度各向异性的测试噪声下也成立。当两类协方差均为对角且足够接近时,行归一化对Adam的优势可能反转,但若对两者施加相同的随机旋转,可通过改变其与Adam坐标轴的对齐方式恢复该优势。合成实验与语言模型最后一层实验均支持上述预测的优势。

英文摘要

Different optimizers can fit the same training data while selecting classifiers with substantially different geometries, but whether this difference provably affects population performance remains unclear. We show that row-wise normalization can achieve strictly higher population accuracy than full-batch Adam, a proxy for random-reshuffling Adam, and exact-SVD Muon in high-dimensional multiclass classification. Under an isotropic Gaussian-cloud data model, this advantage arises because row normalization's class-wise Euclidean geometry asymptotically preserves the population decision-boundary directions, whereas Adam's coordinate-wise geometry and Muon's spectral geometry introduce nonvanishing distortions. Beyond isotropy, the advantage persists for full-batch training on class means with independently oriented class-mean and test-noise covariances. It holds for power-law spectra with class-mean exponent below one, even under heavily anisotropic test noise. When both covariances are diagonal and sufficiently close, the advantage over Adam can reverse, while applying the same random rotation to both restores it by changing only their alignment with Adam's coordinate axes. Synthetic and last-layer language-model experiments support the predicted advantage.

Comments90 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑