arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FedWeave:重新思考异构联邦MoE-LoRA中的专业化单元

FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

Donghang Duan, Xu Zheng, Lizong Zhang, Chong Mu, Meng Han

arXiv 2607.26618首次发表:更新:

AI 中文总结

FedWeave针对异构联邦MoE-LoRA中客户端粒度专业化的局限,提出非对称聚合框架,分离专家聚合与路由优化,在异构多任务基准上优于强基线,实现高效联邦大模型适配。

AI 中文摘要

联邦参数高效微调(PEFT)使大语言模型(LLM)能够协同适配去中心化的私有数据,而无需共享原始样本。然而,客户端之间的任务异质性会在聚合过程中引发跨任务干扰和梯度冲突。联邦MoE-LoRA通过专业化的LoRA专家和条件路由解决了这一挑战,但现有方法通常在客户端粒度进行专业化,隐含假设客户端任务一致。我们的核心见解是:专家需要纯度,即能保持专业化的模式一致更新;而路由需要对比,即支持专家比较的混合任务观测。我们提出FedWeave,一个采用非对称聚合的框架,将专家聚合与路由优化分离以满足这两个需求。FedWeave使用无监督原型发现形成本地桶并在客户端间对齐,实现原型级专家聚合,同时保留混合任务客户端轨迹用于路由训练。推理时,FedWeave使用一个活跃专家执行稀疏推理,同时保留几乎所有软路由性能。我们的理论分析解释了非对称聚合的优势:它通过非模式污染控制专家在平稳性中的收敛,识别由碎片化路由轨迹导致的共识误差,并界定稀疏推理风险。在使用主流LLM主干的异构多任务基准上,FedWeave始终优于强基线, ablation研究验证了我们设计的有效性。

英文摘要

Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation. Federated MoE-LoRA addresses this challenge through specialized LoRA experts and conditional routing. Yet existing methods typically specialize at client granularity, implicitly assuming task-coherent clients. Our core insight is that experts need purity, namely pattern-coherent updates that preserve specialization, whereas routers need contrast, namely mixed-task observations that support expert comparison. We propose FedWeave, a framework that adopts asymmetric aggregation, separating expert aggregation from router optimization to meet these two requirements. FedWeave uses unsupervised prototype discovery to form local buckets and align them across clients, enabling prototype-level expert aggregation while retaining mixed-task client trajectories for router training. At inference, FedWeave performs sparse inference with one active expert while preserving nearly all soft-routing performance. Our theoretical analysis explains why asymmetric aggregation is advantageous: it controls expert convergence in stationarity through off-pattern contamination, identifies the consensus error induced by fragmented router trajectories, and bounds sparse-inference risk. On a heterogeneous multi-task benchmark with mainstream LLM backbones, FedWeave consistently outperforms strong baselines, while ablations verify the effectiveness of our design.

Comments14 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑