arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

联邦轻量级微调

Federated Lightweight Fine-Tuning

Radhakrishna Achanta, Will Reed

arXiv 2607.18343首次发表:更新:

发表机构

Cisco Systems Inc.(思科系统公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究联邦微调通信瓶颈问题,提出基于映射网络的方法,通过低秩分解和增量公式等改进,实现低带宽通信,在CIFAR - 网络上大幅减少通信量,准确率与全权重FedAvg相近,在带宽 - 准确率方面表现优异。

AI 中文摘要

联邦微调受通信瓶颈限制:FedAvg和伪梯度方案传输的负载随模型规模增长,梯度压缩仅能以常数因子缩小负载。我们采用不同方法,映射网络通过冻结的仿射投影从小的可训练潜变量生成网络权重,因映射是共享且仿射的,平均潜变量即平均生成的权重。通过两项改变将其转化为实用的低带宽联邦通道:投影的低秩、种子可再生分解,以及围绕共享的中心预训练基础学习加法校正的增量公式。冻结的正交分类器头进一步提高准确性并减少负载。在CIFAR - 100数据集上使用ResNet - 18 + GroupNorm,我们的方法(FLITE)每轮每个客户端通信1280个浮点数,减少了8718倍,准确率达到74.67%,与全权重FedAvg相差约0.5个百分点。平均恒等式在浮点精度内成立,该方法在带宽 - 准确率帕累托图上比PowerSGD和top - k低一到两个数量级,在强非IID偏斜下匹配或超过全权重FedAvg。使用int4潜变量时,每轮达到648字节且准确率不变,而int4全权重FedAvg则降至随机水平。

英文摘要

Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor. We take a different lever. Mapping networks generate a network's weights from a small trainable latent through a frozen affine projection; because the map is shared and affine, averaging latents is exactly averaging the generated weights. We turn this into a practical low-bandwidth federated channel with two changes: a low-rank, seed-regenerable factorisation of the projection (cutting generator memory from ~80 GB to ~10 MB), and a delta formulation $θ= θ^{\mathrm{pre}} + U V^{\top} z$ that learns an additive correction around a shared centrally-pretrained base -- federated fine-tuning, which is what makes the method work at scale. A frozen orthogonal classifier head further removes the head from the payload while improving accuracy. On CIFAR-100 with ResNet-18+GroupNorm, our method (FLITE, Federated Low-rank Iterative Training Engine) communicates 1,280 floats (~5 KB) per client per round -- an 8718x reduction -- and reaches 74.67%, within ~0.5 pp of full-weight FedAvg. The averaging identity holds to floating-point precision ($6 \times 10^{-8}$); the method sits one to two orders of magnitude below PowerSGD and top-k on the bandwidth-accuracy Pareto; it matches or exceeds full-weight FedAvg under strong non-IID skew. int4 latents reach 648 bytes per round at unchanged accuracy, whereas int4 full-weight FedAvg collapses to chance.

Comments22 pages, 10 figures, 6 tables. Extended preprint with appendix. Under review at ACCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑