arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03941cs.LG

缪昂优化器遇上Mamba:状态空间模型的谱优化

Muon Meets Mamba: Spectral Optimization for State Space Models

Arslan Battalov, Karim Kramin, Alexander Markotenko, Sofia Sinitsina

首次发表
浏览论文内容

中文总结 AI 辅助

本文对比缪昂与AdamW在Mamba-2 1.3亿参数模型上的表现,发现仅在输出投影使用缪昂可提升令牌效率,该优势与条件数无关。

中文摘要 AI 辅助

缪昂(Muon)是一种新型优化器,它通过牛顿-舒尔茨迭代将每个权重矩阵的更新正交化,该迭代在谱范数下执行最速下降。目前几乎所有关于它的证据都来自Transformer模型,其在状态空间模型上的表现尚未有较多报道。我们在仅改变哪些权重组使用缪昂训练的受控协议下,将缪昂与AdamW在Mamba-2 1.3亿参数模型上进行对比。缪昂的优势具有局部性:仅在输出投影上使用缪昂,其效果优于在输入投影或同时在两者上使用缪昂。该优势主要体现在令牌效率上,在两个语料库、两个令牌预算下均成立,且在训练远超计算最优点时仍持续存在。条件数无法解释该增益:缪昂会降低其训练的任意投影的条件数,但条件数更好的输入投影并非带来帮助的那个。

英文摘要

Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. We compare Muon with AdamW on Mamba-2 130M under a controlled protocol that varies only which weight groups are trained with Muon. The benefit is localized. Muon on the output projection alone beats Muon on the input projection or on both. The advantage is mainly one of token efficiency. It holds on two corpora and two token budgets, and persists when training continues well past the compute-optimal point. Conditioning does not explain the gain. Muon lowers the condition number of whichever projection it trains, but the better-conditioned input projection is not the one that helps.

补充信息

↑