发表机构
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对联邦LoRA中DP聚合不匹配和二次噪声放大问题,提出FedHSIP框架,通过低维参数化将双线性聚合转为线性,并引入异质性感知等距投影,在隐私保护下提升性能3-4%并降低80%通信成本。
AI 中文摘要
联邦低秩自适应(LoRA)为在分布式和隐私敏感数据上微调大型语言模型提供了一种高效解决方案。然而,尽管避免了原始数据共享,联邦LoRA仍然容易通过传输的模型更新导致隐私泄露。差分隐私(DP)可以缓解此类泄露,但将DP集成到联邦LoRA中引入了两个基本挑战:独立平均低秩因子导致的聚合不匹配,以及当噪声注入两个因子时产生的二次噪声放大。为了解决这些挑战,我们提出了FedHSIP,一个基于统一低维参数化的差分隐私联邦LoRA框架。FedHSIP将所有LoRA参数重构为一个共享的低维可训练向量,使客户端只需优化和通信低维更新。这种重构将联邦LoRA从双线性因子聚合问题转化为统一的线性参数空间,从而消除了聚合不匹配并防止了DP噪声的二次放大。为了进一步处理非独立同分布(non-IID)数据,我们引入了一种基于预热统计构建的异质性和敏感性感知的等距投影,该投影将具有兼容跨客户端更新模式的坐标分组,同时在低维空间中平衡敏感性、更新能量和异质性。在自然语言理解和生成基准上的大量实验表明,FedHSIP在私人和非私人设置下均持续优于现有的联邦LoRA方法,在差分隐私下实现了高达3-4%的提升,同时将通信成本降低了超过80%,并在异构数据分布下保持了鲁棒性。
英文摘要
Federated Low-Rank Adaptation (LoRA) provides an efficient solution for finetuning large language models across distributed and privacy-sensitive data. However, despite avoiding raw data sharing, federated LoRA remains vulnerable to privacy leakage through transmitted model updates. Differential privacy (DP) mitigates such leakage, but integrating DP into federated LoRA introduces two fundamental challenges: aggregation mismatch from independently averaging low-rank factors, and quadratic noise amplification when noise is injected into both factors. To address these challenges, we propose FedHSIP, a differentially private federated LoRA framework based on a unified low-dimensional parameterization. FedHSIP reformulates all LoRA parameters into a shared low-dimensional trainable vector, enabling clients to optimize and communicate only low-dimensional updates. This reformulation transforms federated LoRA from a bilinear factor aggregation problem into a unified linear parameter space, thereby eliminating aggregation mismatch and preventing the quadratic amplification of DP noise. To further handle non-IID data, we introduce a heterogeneity- and sensitivity-aware isometric projection, constructed from warm-up statistics, which groups coordinates with compatible cross-client update patterns while balancing sensitivity, update energy, and heterogeneity across the low-dimensional space. Extensive experiments on natural language understanding and generation benchmarks show that FedHSIP consistently outperforms existing federated LoRA methods under both private and non-private settings, achieving up to 3-4% improvements under differential privacy while reducing communication cost by over 80% and maintaining robustness under heterogeneous data distributions.