arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dion3:全栈正交更新

Dion3: Full-Stack Orthogonal Updates

Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford

arXiv 2608.11612首次发表:更新:

发表机构

Microsoft Research; New York University; Princeton University; NVIDIA; Yale University(微软研究院; 纽约大学; 普林斯顿大学; 英伟达; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Dion3是针对Muon优化器开销的改进版本,通过全栈优化降低了正交化、通信等开销,步长时间最多降6倍,可作为Muon的即插即用替代方案。

AI 中文摘要

Muon优化器因具有三次时间复杂度的Newton-Schulz正交化步骤产生显著开销,当权重分片时,通信开销会加剧该计算开销,在诸多场景中削弱Muon的优势。我们提出Dion3,这是针对Muon在全栈各层级开销进行优化的改进版本。我们的Gram Newton-Schulz算法降低了正交化的浮点运算(FLOP)开销,CuteDSL内核通过利用对称性加速正交化,megabatching策略减少了通信开销。此外,我们对更新规则提出简单修改以进一步降低成本:每步仅选择动量矩阵的部分行进行正交化。该更新规则在速度和性能上均优于Dion(Muon的另一种“压缩”版本)。总体而言,Dion3达到或优于Muon的损失值,但优化器步长时间最多降低6倍。Dion3可通过dion包获取,作为Muon的即插即用替代方案。

英文摘要

The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. We present Dion3, a revision of Muon that targets this overhead at every level of the stack. Our Gram Newton-Schulz algorithm reduces the FLOP cost of orthogonalization, our CuteDSL kernels accelerate it by exploiting symmetry, and our megabatching strategy reduces communication overhead. Moreover, we propose a simple change to the update rule that cuts costs even further: selecting only a fraction of the momentum matrix's rows to orthogonalize at each step. This update rule improves on Dion (another "compressed" version of Muon), in both speed and performance. Overall, Dion3 matches or improves on the loss achieved by Muon but reduces optimizer step time by up to 6x. Dion3 is available via the dion package (https://github.com/microsoft/dion) as a drop-in replacement for Muon.

Comments37 pages, 23 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑