发表机构
Tianjin University(天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FedLore通过每轮共享并跨轮刷新的低秩投影基,解决联邦学习中子空间碎片化导致的聚合偏差,实现通信和内存高效,并在视觉与语言任务上达到或超越全参数训练。
AI 中文摘要
基础模型的联邦训练受到客户端内存和通信成本的制约。基于LoRA的方法通过低秩适配器降低了这些成本,但其固定的秩预算可能限制适配能力。梯度低秩优化提供了更大的灵活性,然而独立选择的客户端子空间产生了一个我们称之为“子空间碎片化”的问题:局部投影与数据异质性相互作用,使聚合方向产生偏差,同时聚合可能增加更新秩和通信成本。因此,准确的局部梯度压缩未必能保持全局下降。我们提出了FedLore,它在每一轮内共享一个低秩优化基,并在各轮之间刷新该基。共享基使得在低秩坐标下能够进行精确聚合,并消除了所识别的投影偏差。子空间刷新允许累积的模型更新超过每轮的秩预算。我们刻画了聚合偏差,并在全局梯度覆盖条件、标准光滑性和方差假设以及有界梯度异质性下,为投影SGD变体建立了O(T^{-1/2})的平稳性界。在视觉和语言任务上的实验(包括联邦预训练)表明,FedLore优于所评估的低秩适配器基线,并达到或超过全参数训练,同时减少了通信和优化器状态内存。
英文摘要
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.