发表机构
Khalifa University; Central South University; Université de Lorraine(哈利法大学; 中南大学; 洛林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对联邦微调LLM的通信开销与LoRA聚合困境,提出FedFit,采用不相交共享向量库参数化与交替优化,结合残差谱聚合和块级量化,实现高达100倍压缩且性能相当。
AI 中文摘要
联邦学习(FL)能够实现大语言模型(LLMs)的隐私保护微调,但巨大的通信开销仍是关键瓶颈。此外,在联邦学习中应用低秩适配(LoRA)面临基本“聚合困境”,即在精确的乘积和(SoP)与通信高效的乘积和(PoS)实现之间的权衡。为应对这些挑战,我们提出FedFit。首先,为显著降低通信开销,我们引入一种不相交共享向量库参数化方法,从两个紧凑且不相交的全局向量库重建高维适配器矩阵。其次,为解决聚合困境,我们设计了一种交替优化调度。通过在解耦的单库更新(允许精确聚合)与由残差谱聚合机制校正的联合更新之间循环,我们解决了SoP与PoS之间的冲突。此外,我们整合了块级量化与客户端误差反馈,以进一步压缩传输的向量。更进一步,我们为所提算法建立了理论收敛保证。在Qwen2.5模型上的大量实验表明,FedFit实现了与标准联邦LoRA方法相当的困惑度性能,同时提供高达100倍的压缩比。
英文摘要
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.