AI 中文总结
研究针对联邦微调中资源限制、系统异构性和非IID数据等挑战,提出统一框架FeDiSyn。通过缩放定律、基于扩散的合成图像生成及贡献感知LoRA配置和带宽分配算法,减少训练时间和通信成本,提升训练和准确率。
AI 中文摘要
联邦微调(FedFT)能够在分布式、对隐私敏感的设备上调整预训练的大型视觉模型(LVM),但其实际部署面临资源限制、系统异构性和非IID数据这三个关键挑战。先前研究虽部分解决了这些问题,但仍不充分且零散。具体而言,现有合成图像生成方法无法捕捉特定设备的特征分布,基于参数高效微调(PEFT)的FedFT方法往往忽视可能提供关键信息的较弱设备。更重要的是,预训练和FedFT的单独优化忽略了它们的内在联系。为克服这些限制,我们提出了FeDiSyn,一个统一框架,全面考虑预训练和FedFT之间的相互作用,以最小化整体LVM训练时间。具体来说,FeDiSyn引入了:(i)用于FedFT预训练的缩放定律,以确定合成图像的最佳数量,平衡预训练收益与生成/预训练成本;(ii)基于扩散的合成图像生成,捕捉特定设备的特征分布用于预训练,以解决非IID数据问题;(iii)一种用于FedFT的贡献感知LoRA配置和带宽分配算法,以确保在解决系统异构性的同时有效利用信息丰富的设备。在实际测试平台上的实验结果表明,FeDiSyn将完成时间减少了超过52.5%,通信成本减少了超过97.2%,同时实现了与现有技术解决方案相当的准确率。
英文摘要
Federated fine-tuning (FedFT) enables adapting pre-trained large vision models (LVMs) on distributed, privacy-sensitive devices, while its practical deployment is hindered by three critical challenges: resource constraints, system heterogeneity, and non-IID data. While prior studies partially address these issues, e.g., by pre-training initial models on synthetic images to mitigate the adverse effects of non-IID data, or leveraging parameter-efficient fine-tuning (PEFT) methods like low-rank adaptation (LoRA) to reduce resource consumption, they remain inadequate and fragmented. Specifically, existing synthetic image generation methods fail to capture device-specific feature distributions, while current PEFT-based FedFT methods often undervalue weaker devices that may provide critical information. More importantly, the separate optimization of pre-training and FedFT neglects their inherent connection, lacking a holistic perspective to maximize training efficiency. To overcome these limitations, we propose FeDiSyn, a unified framework that holistically considers the interplay between pre-training and FedFT to minimize the overall LVM training time. Specifically, FeDiSyn introduces: (i) a scaling law for FedFT pre-training to determine the optimal number of synthetic images, balancing pre-training benefit against generation/pre-training cost, (ii) diffusion-based synthetic image generation that captures device-specific feature distributions for pre-training to tackle non-IID data, and (iii) a contribution-aware LoRA configuration and bandwidth allocation algorithm for FedFT to ensure that informative devices are effectively utilized while addressing system heterogeneity. Experimental results on the real-world testbed demonstrate that FeDiSyn reduces completion time by over 52.5% and communication cost by over 97.2%, while achieving comparable accuracy to state-of-the-art solutions.
CommentsThis paper has been accepted by International Conference on Parallel Processing (ICPP 2026)