arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AF-Muon:用于权重共享嵌入模型的免AdamW Muon优化器

AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models

Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu

arXiv 2610.01395首次发表:更新:

发表机构

LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University; National University of Singapore(上海交通大学人工智能学院LUMIA实验室; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AF-Muon提出免AdamW的Muon扩展,针对权重共享词汇表采用有限容量线性最小化预言机,在九种设置中优于混合Muon和Sign,节省约20%状态内存。

AI 中文摘要

Muon通过将谱范数最陡下降更新应用于矩阵参数来改善大规模训练,但实际模型也包含不适合稠密矩阵几何的参数块。一个重要的例子是权重共享的词汇表,它出现在语言模型和其他令牌生成器中,可以接收多种结构不同的梯度来源,从稀疏的输入查找到稠密的输出分类器更新。在参考方案中,这些块被交给辅助的AdamW优化器,该优化器恢复二阶矩状态并将别名表作为通用张量更新。我们提出AF-Muon,这是Muon的一种免AdamW扩展,它保持Muon对隐藏权重矩阵的矩阵更新,同时使用支持感知的有限容量线性最小化预言机来处理权重共享词汇表,并使用RMS归一化更新来处理一维辅助参数。因此,AF-Muon用单个一阶矩缓冲区训练每个参数类别,且没有二阶矩状态,在我们的基准测试中,相对于混合Muon,节省了约20%的优化器状态内存。在九种权重共享令牌设置中——从124M到1B参数的仅解码器语言模型、完全共享的T5风格编码器-解码器,以及ImageGPT风格的图像令牌、蛋白质和稀疏MoE变体,涵盖文本、图像和蛋白质序列数据——AF-Muon在平均验证损失和困惑度上均优于混合Muon和SCION风格的Sign端点。长期运行和超参数敏感性研究证实了这种增益的稳健性,相同的动量诊断将其归因于有限容量,该容量比Sign保留了更多的行内幅度,同时限制了行RMS的坐标集中度。这些结果将权重共享词汇表识别为一种独特的优化器几何结构,并在模型、模态和架构中产生了一种稳健的免AdamW Muon变体,在匹配训练中仅增加约1%的步时开销。

英文摘要

Muon improves large-scale training by applying a spectral-norm steepest-descent update to matrix parameters, but practical models also contain parameter blocks that do not fit dense-matrix geometry. One important case is the tied vocabulary table, which appears in language models and other token generators and can receive multiple structurally different gradient sources, from sparse input lookups to dense output-classifier updates. In the reference recipe these blocks are handed to an auxiliary AdamW optimizer, which restores second-moment state and updates the aliased table as a generic tensor. We propose AF-Muon, an AdamW-free extension of Muon that keeps the Muon matrix update for hidden weight matrices while using a support-aware finite-cap linear minimization oracle for tied vocabulary tables and an RMS-normalized update for one-dimensional auxiliary parameters. AF-Muon therefore trains every parameter class with a single first-moment buffer and no second-moment state, saving around 20% optimizer-state memory relative to Hybrid Muon in our benchmark. Across nine tied-token settings - decoder-only language models from 124M to 1B parameters, a fully shared T5-style encoder-decoder, and ImageGPT-style image-token, protein, and sparse-MoE variants, spanning text, image, and protein-sequence data - AF-Muon improves mean validation loss and perplexity over both Hybrid Muon and a SCION-style Sign endpoint. Long-horizon runs and hyperparameter sensitivity studies confirm the gain is robust, and identical-momentum diagnostics attribute it to the finite cap, which preserves more within-row magnitude than Sign while bounding the coordinate concentration of row-RMS. These results identify tied vocabulary tables as a distinct optimizer geometry and yield a robust AdamW-free Muon variant across models, modalities, and architectures, with about 1% step-time overhead in matched training.

Comments63 pages, 19 figures. An earlier, shorter version of this work was accepted as a poster at the OPT 2026 workshop (Optimization for Machine Learning) at NeurIPS 2026; this is the complete version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑