arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05895cs.DC

MatrixFSDP:基于ZeRO-3参数分片的无通信矩阵优化器

MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding

Ming Gao, Yanwu Xu, Hao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对矩阵优化器与ZeRO-3参数分片的系统不匹配问题。核心方法是改变ZeRO-3分片位置,通过多种技术实现无通信矩阵优化。主要贡献是在保持内存优势下,显著降低优化器步骤延迟,实现端到端加速,支持更大模型运行。

中文摘要 AI 辅助

诸如Muon的矩阵优化器对大规模训练具有吸引力,因其能提升收敛性和令牌效率。Muon通过牛顿 - 舒尔茨方法对动量平滑矩阵更新进行正交化,产生频谱平衡更新,需完整二维矩阵作为输入。这导致系统不匹配,现有系统要么在每次优化器步骤重建矩阵,要么使用ZeRO-1所有者放置使更新局部化。MatrixFSDP采取第三条路径,改变ZeRO-3分片位置,一个数据并行等级拥有整个矩阵,其他等级持有空分片,非矩阵张量打包到尾部所有者并保留在AdamW上。普通反向归约将完整Muon输入落在所有者上,牛顿 - 舒尔茨在本地运行,无需优化器步骤矩阵集合。通过MatrixShard元数据等方法,MatrixFSDP在保持ZeRO-3规模内存的同时,更新与全矩阵Muon匹配,在64个A100上,相比库存FSDP2-Muon,单节点优化器步骤延迟降低4.2倍,八节点降低54.6倍,实现高达2.15倍的端到端加速,且能运行ZeRO-1所有者放置超过80GB GPU的模型大小。

英文摘要

Matrix optimizers such as Muon are attractive for large-scale training because they can improve convergence and token efficiency over coordinate-wise optimizers. Muon does this by orthogonalizing momentum-smoothed matrix updates with Newton-Schulz, producing spectrum-balanced updates that require the complete 2D matrix as input. This exposes a systems mismatch: FSDP/ZeRO-3 saves memory by making the optimizer see shards, not whole matrices. Existing systems therefore either reconstruct matrices at every optimizer step, paying weight-sized communication after backward, or make the update local by using ZeRO-1 owner placement with full parameters resident. MatrixFSDP takes a third path: it changes where ZeRO-3 shards live, not the optimizer being computed. For each 2D weight, one data-parallel rank owns the whole matrix and the other ranks hold empty shards; non-matrix tensors are packed into tail owners and stay on AdamW. The ordinary backward reduction then lands the full Muon input on the owner, so Newton-Schulz runs locally with no optimizer-step matrix collective. Forward and backward still materialize and reshard parameters; the runtime challenge is to make that uneven layout efficient and correct. MatrixFSDP does so with MatrixShard metadata, a balance-aware owner planner, deterministic owner-segment P2P collectives, owner-buffer pinning, and owner-shard checkpoint resharding. The resulting update matches full-matrix Muon while preserving ZeRO-3-scale memory: on 64 A100s, MatrixFSDP reduces optimizer-step latency over stock FSDP2-Muon by 4.2x on one node and 54.6x on eight nodes, reaches up to 2.15x end-to-end speedup, and runs model sizes where ZeRO-1 owner placement exceeds an 80 GB GPU.

发表机构

  • University of Pittsburgh(匹兹堡大学)
  • Google(谷歌)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑