发表机构
Georgia State University(佐治亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE大语言模型联邦微调中异构客户端数据导致的专家专业化削弱和更新冲突问题,提出任务感知方法FedTAR,利用路由输出和SVD提取任务坐标进行分层聚合,保持专家专业化并取得最优性能。
AI 中文摘要
混合专家(Mixture-of-Experts, MoE)已成为大语言模型(LLMs)广泛采用的架构,因为它通过稀疏专家激活在提升模型容量的同时限制了计算开销。这一特性使得基于MoE的大语言模型在资源受限的分布式环境中特别具有吸引力。然而,在异构客户端数据下,基于MoE的大语言模型的联邦微调仍然具有挑战性。由于客户端通常对应不同的任务偏好,直接聚合其本地更新可能会削弱专家专业化,并在共享专家上引入冲突的更新方向。为应对这些挑战,我们提出了FedTAR,一种针对基于MoE的大语言模型的任务感知联邦微调方法。FedTAR通过路由输出建立本地更新与任务偏好之间的关联。具体而言,我们将奇异值分解(SVD)应用于路由特征和本地更新,以提取低维任务坐标和更新方向。基于任务坐标,FedTAR在具有相似任务偏好的客户端之间执行簇内聚合,并在不同任务组之间执行簇间聚合。然后通过学到的任务到更新的映射重建聚合更新,确保最终更新与任务特定的优化方向保持一致。通过这种方式,FedTAR保持了专家专业化并减轻了异构客户端之间的破坏性干扰。我们在四种不同的非独立同分布(non-IID)设置下的基准任务上评估了FedTAR。实验结果表明,FedTAR始终优于强联邦微调基线,并实现了最先进的性能。
英文摘要
Mixture-of-Experts (MoE) has become a widely adopted architecture for Large Language Models (LLMs), as it improves model capacity while limiting computational overhead through sparse expert activation. This property makes MoE-based LLMs particularly attractive for resource-constrained distributed environments. However, federated fine-tuning of MoE-based LLMs remains challenging under heterogeneous client data. Since clients often correspond to different task preferences, directly aggregating their local updates may weaken expert specialization and introduce conflicting update directions on shared experts. To address these challenges, we propose FedTAR, a task-aware federated fine-tuning method for MoE-based LLMs. FedTAR establishes the association between local updates and task preference via routing outputs. Specifically, we apply Singular Value Decomposition (SVD) to both routing features and local updates to extract low-dimensional task coordinates and update directions. Based on the task coordinates, FedTAR performs intra-cluster aggregation among clients with similar task preferences and inter-cluster aggregation across different task groups. The aggregated update is then reconstructed through the learned task-to-update mapping, ensuring that the final update remains aligned with task-specific optimization directions. In this way, FedTAR preserves expert specialization and mitigates destructive interference among heterogeneous clients. We evaluate FedTAR on four benchmark tasks under different non-IID settings. Experimental results demonstrate that FedTAR consistently outperforms strong federated fine-tuning baselines and achieves state-of-the-art performance.
CommentsAccepted by ICDM2026