发表机构
University of Electronic Science and Technology of China; Fudan University; Singapore University of Technology and Design(电子科技大学; 复旦大学; 新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对联邦LoRA的因子共享问题,提出FedAS-LoRA方法,通过RSS指标在训练前选择共享侧,提升了大语言模型联邦微调的性能。
AI 中文摘要
低秩适配(LoRA)用两个紧凑矩阵因子(即$A$和$B$)表示大语言模型(LLM)的更新,为联邦学习范式下的大模型微调提供了高效方式。受LoRA因子非对称角色的启发,我们研究是让$A$在客户端间共享、$B$保持客户端专属(Share-A/Local-B),还是让$B$共享、$A$保持客户端专属(Share-B/Local-A)。通过最小二乘代理,我们发现Share-A/Local-B要求客户端专属LoRA更新矩阵使用公共秩$r$的输入侧空间,而Share-B/Local-A要求公共秩$r$的输出侧空间。两种策略因此产生不同的投影残差,表明优选策略是客户端间聚合残差更小的那一种。基于这一见解,我们提出联邦自适应因子共享低秩适配(FedAS-LoRA),它在训练前选择共享侧以提升微调性能。为实现训练前的自适应因子选择,我们设计了秩感知共享子空间充分性(RSS)指标,该指标利用冻结的LLM主干提取的表示,有效评估共享秩$r$的输入子空间是否适配本地数据分布。在不同任务、数据分布、LoRA秩和参与设置下的实验证实了RSS的有效性和FedAS-LoRA的优越性能。
英文摘要
Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead be shared while $A$ remains client-specific (Share-B/Local-A). With a least-squares surrogate, we reveal that Share-A/Local-B requires the client-specific LoRA update matrices to use a common rank-$r$ input-side space, whereas Share-B/Local-A requires a common rank-$r$ output-side space. The two strategies therefore incur different projection residuals, indicating that the preferred strategy is the one with the smaller aggregate residual across clients. With this insight, we propose Federated Adaptive Factor Sharing Low-Rank Adaptation (FedAS-LoRA), which selects the sharing side before training to enhance fine-tuning performance. To enable adaptive factor selection before training, we design a Rank-Aware Shared-Subspace Sufficiency (RSS) metric, which effectively assesses whether a shared rank-$r$ input subspace is sufficient for the local data distributions using representations extracted from a frozen LLM backbone. Experiments across different tasks, data distributions, LoRA ranks, and participation settings confirm the effectiveness of RSS and the superior performance of FedAS-LoRA.
Comments20 pages