LionMuon:交替谱下降与符号下降以实现高效训练
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
浏览论文内容
中文总结 AI 辅助
LionMuon通过交替使用Muon谱步骤与Lion符号步骤,每P步执行一次Muon,共享动量缓冲区,在保持低通信开销的同时,以更少的计算和内存实现比Muon、AdamW、Lion和Signum更低的训练损失。
中文摘要 AI 辅助
预训练语言模型需要巨大的计算量,而正确的优化器可以节省其中相当一部分。Muon的谱步骤比符号步骤提供了更强的方向,但代价高昂。每一步都对完整矩阵运行Newton-Schulz迭代,并且在分布式训练中,还需要额外的全归约操作。而Lion和Signum中的符号步骤则成本低廉,且仅在各设备本地进行。我们提出了LionMuon,它每P次迭代执行一次Muon步骤,其间穿插Lion步骤,并共享一个双指数移动平均(EMA)动量缓冲区。Muon的计算和通信开销每P步支付一次,且优化器状态仅为AdamW的一半。单EMA变体SignMuon已经优于Muon。我们在重尾噪声下证明了复杂度界限,其中周期在Muon和Lion的平滑度与噪声常数之间进行插值,并说明了LionMuon何时比两者都快。在FineWeb上训练的124M和355M模型中,使用P=2和P=5的LionMuon在相同token数量下达到了比Muon、AdamW、Lion和Signum更低的损失。在4-GPU数据并行训练中,它在PCIe上以少三分之一的墙钟时间达到了Muon的最终损失,并且在不超过暴露通信的情况下,在损失上优于通信高效的Muon变体Dion和MuonBP,同时保持精确梯度。代码:此https URL
英文摘要
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon