arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FedSubMuon:通过结构化子空间Muon实现通信高效联邦大语言模型微调

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang

arXiv 2609.06073首次发表:更新:

发表机构

King Abdullah University of Science and Technology (KAUST); Imperial College London; Technische Universität Berlin (TU Berlin); Ocean University of China; The Hong Kong Polytechnic University(阿卜杜拉国王科技大学; 帝国理工学院; 柏林工业大学; 中国海洋大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FedSubMuon,在共享结构化子空间内优化系数矩阵,实现通信高效的联邦Muon微调,并扩展FedSubMuon-GT提升精度,实验证明其在多数任务上精度最优且通信成本显著降低。

AI 中文摘要

联邦微调将大型语言模型(LLM)适应到分散的客户端数据,但其在跨设备训练中的可扩展性常常受到高通信成本的限制。Muon是一种优化器,通过为矩阵值参数正交化动量来提升优化性能。现有的联邦Muon方法展示了矩阵感知优化在联邦学习中的优势,但仍需传输完整的层大小更新和优化器状态。减少通信的一种自然方式是直接将Muon应用于LoRA因子,但这会改变优化目标并削弱Muon的矩阵感知更新几何。我们提出FedSubMuon,一种通信高效的联邦Muon微调方法,它在共享的结构化子空间内优化紧凑的系数矩阵。这种设计使Muon保持作用于单个矩阵值可训练对象,同时将客户端上传减少为紧凑的系数矩阵。我们进一步引入FedSubMuon-GT,一种面向精度的扩展,使用投影梯度使跟踪的子空间基适应任务相关的梯度方向。在指令微调和数学推理上的实验表明,FedSubMuon-GT在五个数据集-模型对中的四个上取得了最佳整体精度,而FedSubMuon在所有匹配的通信预算下表现最佳。在Dolly-15K上,最接近的通信基线在Llama-1B和Qwen-4B上分别需要5.5倍和1.4倍的总通信量。

英文摘要

Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces. This design keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices. We further introduce FedSubMuon-GT, an accuracy-oriented extension that uses projected gradients to adapt tracked subspace bases toward task-relevant gradient directions. Experiments on instruction tuning and mathematical reasoning show that FedSubMuon-GT achieves the best overall accuracy on four of five dataset-model pairs, while FedSubMuon performs best under all matched communication budgets. On Dolly-15K, the closest communication baseline requires 5.5 times and 1.4 times more total communication on Llama-1B and Qwen-4B, respectively.

Comments15 pages, 4 figures, 2 algorithms

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑