arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dysco:动态子空间增强以减轻联邦学习中的LoRA干扰

Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning

Haobo Zhang, Jiankun Wang, Suraj Rajendran, Weishen Pan, Lam Tsoi, Yong Chen, Fei Wang, Jiayu Zhou

arXiv 2607.14367首次发表:更新:

发表机构

University of Michigan; Cornell University; University of Pennsylvania(密歇根大学; 康奈尔大学; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对联邦学习中异构客户端使LoRA聚合不稳定的问题,提出动态子空间增强方法Dysco,通过联邦动态分配特定客户端LoRA子空间,经实验验证其能减少干扰、降低训练损失、提升算法性能且开销小。

AI 中文摘要

大型预训练模型的联邦微调越来越依赖低秩适应(LoRA)来减少通信和计算,但异构客户端会使适配器聚合不稳定。我们将数据参数干扰识别为这种不稳定性的几何根源。这种干扰由LoRA更新子空间与客户端激活之间的对齐控制,这表明联邦LoRA聚合不仅应视为参数平均,还应视为子空间分配。我们提出了动态子空间增强(Dysco),这是一种插件方法,以联邦和动态的方式分配特定于客户端的LoRA子空间。在每一轮中,客户端从本地表示中计算对激活不敏感的子空间,并仅传输所得的基;服务器然后通过一个封闭形式的解决方案构建特定于客户端的合并子空间,该解决方案最大化与其他客户端不敏感方向的兼容性。为了处理表示漂移,Dysco执行多轮子空间增强,以保留过去的更新方向,同时适应未来的表示。我们提供了一种收敛分析,将数据参数干扰作为聚合误差项嵌入标准联邦优化界,并证明Dysco的服务器固定合并子空间对该误差产生更紧的上界。在受控的合成联邦任务以及使用Llama-3.2-1B进行的MIMIC-IV临床笔记分类实验表明,Dysco大大减少了干扰,相对于理论确定的正交子空间划分下的基线,最终轮合成训练损失降低了多达9倍,在MIMIC上对所有五种测试的联邦学习算法提高了多达4.3%,优于最近的联邦LoRA方法,并且仅增加了0.9%的挂钟开销。我们的代码可在此https URL上获取。

英文摘要

Federated fine-tuning of large pre-trained models increasingly relies on Low-Rank Adaptation (LoRA) to reduce communication and computation, but heterogeneous clients can make adapter aggregation unstable. We identify the data-parameter interference as a geometric source of this instability. This interference is controlled by the alignment between LoRA update subspaces and client activations, suggesting that federated LoRA aggregation should be viewed not only as parameter averaging but also as subspace allocation. We propose Dynamic Subspace Boosting (Dysco), a plug-in method that allocates client-specific LoRA subspaces in a federated and dynamic manner. In each round, clients compute activation-insensitive subspaces from local representations and transmit only the resulting bases; the server then constructs client-specific merged subspaces through a closed-form solution that maximizes compatibility with other clients' insensitive directions. To handle representation drift, Dysco performs multi-round subspace boosting to preserve past update directions while adapting to future representations. We provide a convergence analysis that embeds the data-parameter interference as an aggregation-error term in a standard federated optimization bound, and prove that Dysco's server-fixed merged subspaces yield a tighter upper bound on this error. Experiments on controlled synthetic federated tasks and on MIMIC-IV clinical-note classification with Llama-3.2-1B show that Dysco substantially reduces interference, reduces the final-round synthetic training loss by up to 9 times relative to baselines under the orthogonal-subspace partition the theory identifies, improves all five tested FL algorithms by up to 4.3% on MIMIC, outperforms recent federated LoRA methods, and adds only 0.9% wall-clock overhead. Our code is available at https://github.com/illidanlab/Dysco.

Comments33 pages, 10 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑