arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MuonIO:用于嵌入表和语言模型头的原则性范数感知下降

MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads

Linkai Ma, Xinyu Luo, Mengbo Wang, Ananth Grama, Petros Drineas, Brian Bullins

arXiv 2610.02705首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MuonIO提出针对嵌入表和语言模型头的统一Muon风格更新,利用范数感知几何,在1B LLaMA预训练中减少50%优化器内存和约46%更新FLOPs,并提升验证困惑度。

AI 中文摘要

Muon优化器通过求解受谱范数惩罚的损失的局部线性化来推导其隐藏线性层的更新规则,其动机是基于对密集线性层的RMS稳定性论证。然而,标准的Muon实现将输入(嵌入表)和输出(语言模型头)层排除在此原则性处理之外,对这些层使用AdamW。我们提出MuonIO,一种针对这两层统一的Muon风格更新。对于语言模型头$\mathbf{L} \in \mathbb{R}^{V \times d}$,由于softmax输出几何的Lipschitz连续性,我们论证使用$2\to\infty$算子范数;而对于嵌入表$\mathbf{E} \in \mathbb{R}^{d \times V}$,我们基于Bernstein & Newhouse (2025)识别出的one-hot输入几何,采用$1 \to 2$算子范数。恒等式$\lVert\mathbf{L}\rVert_{2\to\infty}=\lVert\mathbf{L}^\top\rVert_{1\to2}$将两个矩阵置于相同的词汇导向几何中:MuonIO应用单一归一化向量规则,该规则表现为对$\mathbf{E}$的列归一化和对$\mathbf{L}$的行归一化。实证评估证明了我们方法的有效性,与Muon相比,在C4上进行1B LLaMA预训练时,MuonIO将I/O优化器状态内存减少50%,I/O更新FLOPs减少约46%,同时提高了验证困惑度。

英文摘要

The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (language model head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the language model head $\mathbf{L} \in \mathbb{R}^{V \times d}$, we motivate the use of the $2\to\infty$ operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table $\mathbf{E} \in \mathbb{R}^{d \times V}$, we draw on the $1 \to 2$ operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity $\lVert\mathbf{L}\rVert_{2\to\infty}=\lVert\mathbf{L}^\top\rVert_{1\to2}$ then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for $\mathbf{E}$ and row normalization for $\mathbf{L}$. Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by $\sim$46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑