面向异构联邦指令微调的MoE路由器引导聚类
MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning
另 1 家 · 查看机构详情
- California State University, Dominguez Hills(加州州立大学多明戈斯山分校)
- University of California, Irvine(加州大学欧文分校)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出ClientMorpher框架,通过MoE路由特征设计两种聚类策略,在Databricks Dolly-15K数据集上验证其可提升联邦指令微调的个性化性能,且通信成本与传统方法相当。
中文摘要 AI 辅助
联邦指令微调使大语言模型(LLMs)能在无需共享数据的情况下适配分散且对隐私敏感的数据。近期的混合专家(MoE)大语言模型因稀疏激活可降低计算与通信开销,同时扩展模型容量,在联邦学习中极具吸引力。然而,现有联邦MoE方法主要聚焦参数聚合与个性化,忽视了MoE模型的路由行为可作为客户端协作的信息源。在异构指令分布下,不加区分的聚合会导致负迁移,因此需确定联邦优化过程中哪些客户端应协作。我们提出ClientMorpher,这是一个感知路由的个性化联邦指令微调框架,利用预训练MoE模型的路由特征在聚合前组织客户端协作。我们研究两种互补的聚类策略:ClientMorpher-C直接用专家激活特征对客户端聚类,ClientMorpher-E先基于专家的跨客户端使用特征对专家聚类,再推导客户端协作组。我们在Databricks Dolly-15K数据集上评估ClientMorpher用于联邦指令微调,采用病理分布与基于狄利克雷分布的异构客户端分布,覆盖多个指令跟随任务。实验结果表明,与传统联邦平均和本地训练相比,感知路由的协作在保持相同通信成本的同时,持续提升个性化性能。此外,研究显示以客户端为中心和以专家为中心的聚类,是稀疏MoE大语言模型个性化联邦指令微调的有效且可扩展方法。
英文摘要
Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing. Recent Mixture-of-Experts (MoE) LLMs are particularly attractive for federated learning because their sparse activation reduces computation and communication while scaling model capacity. However, existing federated MoE methods primarily focus on parameter aggregation and personalization, overlooking the routing behavior of MoE models as a source of information for client collaboration. Under heterogeneous instruction distributions, indiscriminate aggregation can lead to negative transfer, highlighting the need to identify which clients should collaborate during federated optimization. We propose ClientMorpher, a routing-aware, personalized federated instruction fine-tuning framework that leverages routing signatures from pretrained MoE models to organize client collaboration prior to aggregation. We investigate two complementary clustering strategies: ClientMorpher-C, which directly clusters clients using expert activation profiles, and ClientMorpher-E, which first clusters experts based on their cross-client usage signatures and then derives client collaboration groups. We evaluate ClientMorpher for federated instruction fine-tuning on the Databricks Dolly-15K dataset, using pathological and Dirichlet-based heterogeneous client distributions across multiple instruction-following tasks. Experimental results show that routing-aware collaboration consistently improves personalized performance compared to conventional federated averaging and local training, while maintaining the same communication cost. Furthermore, our study shows that client-centric and expert-centric clustering provides an effective and scalable approach for personalized federated instruction fine-tuning of sparse MoE LLMs.