发表机构
University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE训练中专家更新与梯度对齐不佳的问题,提出Compass优化器,通过族因子和半径因子调整Muon步长,在FineWeb-Edu预训练中达到或优于现有优化器,并保持专家负载平衡。
AI 中文摘要
混合专家(MoE)语言模型将每个词元发送给少数几个专家,因此每个专家在数据的不同部分上进行训练,并且这部分数据在训练过程中会发生变化。在共享学习率下,Muon 对形状相同的专家矩阵施加大致相同大小的更新,即使专家的更新与其当前梯度对齐不佳时也是如此。我们在此提出 ExpertMuon-Compass(简称 Compass),它将每个专家的 Muon 步长乘以两个因子。一个族因子将专家的正交化更新与其梯度之间的余弦值与该层中其他专家的相同余弦值进行比较。一个标量半径将更新和梯度的对应行之间的对齐程度聚合为一个步长乘数。Compass 保持 Muon 的更新方向和动量缓冲区。在 FineWeb-Edu 上的预训练中,对所有矩阵使用 Nesterov 动量的 Compass 在较长运行中与 Muon、NorMuon 及其他优化器表现相当或更好,且权重衰减与 NorMuon 匹配。将其因子添加到 NorMuon 上可获得与 NorMuon 相同或更低的损失。当每个专家在训练期间看到的数据变化时,Compass 最为有效,例如当多语言语料库的语言以单独块形式到达时。使用 Compass,专家负载保持平衡,路由器更果断地将词元分配给专家。我们还证明了这两个因子的扰动界。
英文摘要
Mixture-of-experts (MoE) language models send each token to a few experts, so each expert is trained on a different part of the data, and this part changes during training. With a shared learning rate, Muon applies updates of roughly the same size to expert matrices of the same shape, even when an expert's update is poorly aligned with its current gradient. We here propose ExpertMuon-Compass (Compass), which multiplies the Muon step of each expert by two factors. A family factor compares the cosine between the expert's orthogonalized update and its gradient with the same cosine for the other experts in its layer. A scalar radius aggregates the alignment between corresponding rows of the update and gradient into one step-length multiplier. Compass keeps the update direction and the momentum buffer of Muon. In pretraining on FineWeb-Edu, Compass with Nesterov momentum on all matrices performs as well as or better than Muon, NorMuon, and other optimizers, with weight decay matched to NorMuon in the longer runs. Adding its factors to NorMuon gives the same or a lower loss than NorMuon. Compass is the most effective when the data seen by each expert varies during training, for example, as when the languages of a multilingual corpus arrive in separate blocks. With Compass, the expert load stays balanced, and the router assigns tokens to experts more decisively. We also prove a perturbation bound for the two factors.
Comments37 pages, 9 figures