arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33147cs.LG

CFLoRA:基于互补因子的联邦大语言模型微调以实现无误差聚合

CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation

Yanan Ma, Qiyuan Chen, Zihan Fang, Xianhao Chen, Yuguang Fang

首次发表
浏览论文内容

中文总结 AI 辅助

CFLoRA通过每轮将LoRA通道划分为互补集合并消除双线性项,实现了无误差的联邦聚合,支持异构秩预算,并达到O(1/√T)收敛速率,在GLUE和常识推理任务上优于现有基线。

中文摘要 AI 辅助

联邦低秩适配(LoRA)使得在不集中私有客户端数据的情况下协作微调大型语言模型成为可能。然而,其因子化更新在联邦平均中造成了结构不匹配:分别平均两个LoRA因子并不等于平均它们的乘积。现有的精确方法主要通过冻结整个因子或跨轮次交替因子来解决这一问题,但没有任何方法能在不产生聚合误差或扩大通信秩的情况下同时更新因子。为解决这一根本问题,我们提出了CFLoRA,一种联邦LoRA方案,它在每一通信轮次中将潜在LoRA通道划分为两个互补集合。通过确保因子间的列和行互补,我们消除了矩阵乘法中的双线性项,使联邦聚合变得精确。关键的是,我们的框架还支持具有异构秩预算的客户端。收敛性分析验证了CFLoRA在同质秩情况下实现了原始LoRA目标的O(1/√T)收敛速率。在GLUE基准上使用RoBERTa以及常识推理任务上使用LLaMA-3.2-3B-Instruct进行的大量实验表明,与最先进的联邦LoRA基线相比,CFLoRA实现了更优的性能和训练效率。

英文摘要

Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their products. Existing exact methods resolve this issue mainly by freezing an entire factor or alternating factors across rounds, but none can update factors simultaneously without aggregation errors or expanding communication ranks. To address this fundamental problem, we present \texttt{CFLoRA}, a federated LoRA scheme that partitions latent LoRA channels into two complementary sets in every communication round. By ensuring that columns and rows are complementary across factors, we eliminate bilinear terms in matrix multiplications, making federated aggregation exact. Crucially, our framework also supports clients with heterogeneous rank budgets. Convergence analysis validates \texttt{CFLoRA} achieves $\mathcal{O}(1/\sqrt{T})$ convergence rate of the \textit{original} LoRA objective in homogeneous-rank cases. Extensive experiments with RoBERTa on the GLUE benchmark and with LLaMA-3.2-3B-Instruct on commonsense reasoning tasks demonstrate that \texttt{CFLoRA} achieves superior performance and training efficiency compared to state-of-the-art federated LoRA baselines.

补充信息

↑