arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展用于大型视觉模型联邦微调的合成图像预训练

Scaling Synthetic-Image Pre-Training for Federated Fine-Tuning of Large Vision Models

Qianpiao Ma, Xiaozhu Song, Junlong Zhou, Yue Zeng, Jianchun Liu, Huaqing Tu

arXiv 2607.12583首次发表:更新:

AI 中文总结

研究针对联邦微调中资源限制、系统异构性和非IID数据等挑战,提出统一框架FeDiSyn。通过缩放定律、基于扩散的合成图像生成及贡献感知LoRA配置和带宽分配算法,减少训练时间和通信成本,提升训练和准确率。

AI 中文摘要

联邦微调(FedFT)能够在分布式、对隐私敏感的设备上调整预训练的大型视觉模型(LVM),但其实际部署面临资源限制、系统异构性和非IID数据这三个关键挑战。先前研究虽部分解决了这些问题,但仍不充分且零散。具体而言,现有合成图像生成方法无法捕捉特定设备的特征分布,基于参数高效微调(PEFT)的FedFT方法往往忽视可能提供关键信息的较弱设备。更重要的是,预训练和FedFT的单独优化忽略了它们的内在联系。为克服这些限制,我们提出了FeDiSyn,一个统一框架,全面考虑预训练和FedFT之间的相互作用,以最小化整体LVM训练时间。具体来说,FeDiSyn引入了:(i)用于FedFT预训练的缩放定律,以确定合成图像的最佳数量,平衡预训练收益与生成/预训练成本;(ii)基于扩散的合成图像生成,捕捉特定设备的特征分布用于预训练,以解决非IID数据问题;(iii)一种用于FedFT的贡献感知LoRA配置和带宽分配算法,以确保在解决系统异构性的同时有效利用信息丰富的设备。在实际测试平台上的实验结果表明,FeDiSyn将完成时间减少了超过52.5%,通信成本减少了超过97.2%,同时实现了与现有技术解决方案相当的准确率。

英文摘要

Federated fine-tuning (FedFT) enables adapting pre-trained large vision models (LVMs) on distributed, privacy-sensitive devices, while its practical deployment is hindered by three critical challenges: resource constraints, system heterogeneity, and non-IID data. While prior studies partially address these issues, e.g., by pre-training initial models on synthetic images to mitigate the adverse effects of non-IID data, or leveraging parameter-efficient fine-tuning (PEFT) methods like low-rank adaptation (LoRA) to reduce resource consumption, they remain inadequate and fragmented. Specifically, existing synthetic image generation methods fail to capture device-specific feature distributions, while current PEFT-based FedFT methods often undervalue weaker devices that may provide critical information. More importantly, the separate optimization of pre-training and FedFT neglects their inherent connection, lacking a holistic perspective to maximize training efficiency. To overcome these limitations, we propose FeDiSyn, a unified framework that holistically considers the interplay between pre-training and FedFT to minimize the overall LVM training time. Specifically, FeDiSyn introduces: (i) a scaling law for FedFT pre-training to determine the optimal number of synthetic images, balancing pre-training benefit against generation/pre-training cost, (ii) diffusion-based synthetic image generation that captures device-specific feature distributions for pre-training to tackle non-IID data, and (iii) a contribution-aware LoRA configuration and bandwidth allocation algorithm for FedFT to ensure that informative devices are effectively utilized while addressing system heterogeneity. Experimental results on the real-world testbed demonstrate that FeDiSyn reduces completion time by over 52.5% and communication cost by over 97.2%, while achieving comparable accuracy to state-of-the-art solutions.

CommentsThis paper has been accepted by International Conference on Parallel Processing (ICPP 2026)

DOI:10.1145/3832810.3832835

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑