面向矩阵型模型的联邦组合Muon优化器
Federated Compositional Muon Optimizer for Matrix-Wise Models
浏览论文内容
中文总结 AI 辅助
针对现有Muon优化器不适用于分层结构化问题的缺陷,提出FedCoMuon及FedCoMuon-VR优化器,理论分析其收敛性并通过实验验证其在联邦学习等任务中的性能优于基线方法。
中文摘要 AI 辅助
Muon是一种近期开发的优化器,适用于AI领域的矩阵型模型。尽管已有诸多研究探讨Muon及其变体,但这些方法仍不太适合分层结构化问题。为填补该空白,本文提出一种有效的联邦组合Muon优化器(FedCoMuon),用于求解分布式矩阵型组合优化问题。具体而言,FedCoMuon优化器基于组合梯度追踪与正交动量构建,还提出基于动量型方差缩减技术的FedCoMuon方差缩减变体(FedCoMuon-VR)。理论上,分析了所提算法在非独立同分布(non-i.i.d.)与非凸设置下的收敛性,尤其证明FedCoMuon-VR在寻找ε-平稳解时,样本复杂度为O(ε⁻³),低于现有FedMuon算法。针对鲁棒联邦学习与任务分布式风险敏感元学习的大量数值实验表明,所提方法与现有组合基线具有竞争力,且在多种设置下达到报告的最佳准确率。
英文摘要
Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. To fill this gap, we propose an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems. Specifically, our FedCoMuon optimizer builds on compositional gradient tracking and orthogonalized momentum. Moreover, we propose a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique. In theory, we analyze the convergence properties of our algorithms under the non-i.i.d. and non-convex settings. In particular, we prove that our FedCoMuon-VR obtains a lower sample complexity of $O(ε^{-3})$ for finding an $ε$-stationary solution than the existing FedMuon algorithms. Extensive numerical experiments on robust federated learning and task-distributed risk-sensitive meta learning show that our proposed methods are competitive with existing compositional baselines and achieve the best reported accuracy in several settings.
发表机构
- College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。