arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视觉Transformer的联邦LoRA中逐层全局秩发现的谱变换

Spectral Transformation for Layer-wise Global Rank Discovery in Federated LoRA for Vision Transformers

Hariharan Ramesh, Jyotikrishna Dass

arXiv 2607.21074首次发表:更新:

发表机构

University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对联邦LoRA微调ViT时现有聚合策略的局限,提出SpecTraL方法,通过谱变换在低秩潜在空间处理本地LoRA模块,利用随机矩阵理论分离信号与噪声,发现全局秩,还引入初始化框架,实验证明该方法提升了精度-通信权衡等性能。

AI 中文摘要

在联邦设置下,使用低秩适配器(LoRA)微调视觉Transformer(ViT)有望提高通信效率,但现有聚合策略存在根本局限性。独立平均LoRA因子在数学上不一致,会引入交叉项聚合误差。通过在服务器上连接本地适配器来保留异构客户端秩的方法会大幅增加下载成本,且常需在客户端将全局LoRA更新合并到预训练权重中,导致重新初始化延迟和收敛不稳定。其他方法通过重建密集权重更新或训练辅助模型来优化聚合误差,进一步增加了服务器端开销。在这项工作中,我们提出了SpecTraL,即用于逐层全局秩发现的谱变换,在统一设计中解决了这些挑战。SpecTraL堆叠来自客户端的本地LoRA模块,并直接在低秩潜在空间中对堆叠的适配器执行正交Householder变换,消除了全局更新的密集重建和服务器上的任何辅助优化。通过利用随机矩阵理论中的尖峰协方差模型,可以从非IID噪声中解析分离全局共识信号,无需手动超参数调整就能发现最优的逐层全局秩。为了在后续轮次中匹配本地秩,我们引入了一个填充感知初始化框架,使客户端能够纳入剩余的LoRA维度,而无需将它们重新合并到预训练的基础模型中。在DomainNet和NICO++上对ViT-B/16和ViT-L/16进行联邦微调的实验表明,在精度-通信权衡方面有所改善,减少了服务器计算,并消除了秩选择的超参数搜索。我们的代码可在这个https URL上获取。

英文摘要

Fine-tuning Vision Transformers (ViTs) with low-rank adapters (LoRA) promises better communication efficiency under federated setup, yet existing aggregation strategies face fundamental limitations. Independently averaging these LoRA factors is mathematically inconsistent, introducing cross-term aggregation error. In contrast, approaches that preserve heterogeneous client ranks by concatenating local adapters on the server substantially increase download cost and often require merging global LoRA updates into pretrained weights on the clients, causing reinitialization lag and unstable convergence. Other approaches further increase server-side overhead by reconstructing dense weight updates or training auxiliary models to refine aggregation error. In this work, we propose SpecTraL, spectral transformation for layer-wise global rank discovery, that resolves these challenges within a unified design. SpecTraL stacks local LoRA modules from clients and performs orthonormal Householder Transformation of the stacked adapters directly in the low-rank latent space, eliminating dense reconstruction of the global update and any auxiliary refinement on the server. By leveraging the Spiked Covariance Model from Random Matrix Theory, SpecTraL analytically separates the global consensus signal from non-IID noise, discovering optimal layer-wise global ranks without manual hyperparameter tuning. To match local ranks in subsequent rounds, we introduce a padding-aware initialization framework that lets clients incorporate residual LoRA dimensions without re-merging them into the pre-trained base model. Experiments on federated fine-tuning of ViT-B/16 and ViT-L/16 over DomainNet and NICO++ demonstrate improved accuracy-communication trade-offs, reduced server computation, and elimination of hyperparameter search for rank selection. Our code is available at https://github.com/DASS-Lab-Group/SpecTraL

CommentsAccepted at ECML-PKDD 2026 (Research Track). This is the submitted version, prior to peer review. Code: https://github.com/DASS-Lab-Group/SpecTraL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑