发表机构
University of Hong Kong; City University of Hong Kong(香港大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出HO-FL混合阶联邦学习框架,通过按内存预算划分模型为ZO和FO段,在异构边缘设备上平衡内存与收敛,实现接近全FO性能并降低内存需求。
AI 中文摘要
在内存受限的边缘设备上进行联邦学习(FL)面临一个两难困境:一阶(FO)优化(即反向传播)需要大量内存,而零阶(ZO)优化则收敛速度严重减慢。为解决这一困境,我们提出了HO-FL,一种混合阶联邦学习框架,该框架使用ZO优化训练模型的底部段,使用FO优化训练其顶部段。每个设备可以根据其内存预算灵活选择其阶数边界,同时参与同一全局模型的训练。此外,我们的收敛性分析揭示了一个新的基本权衡:拥有更大FO训练段的客户端可以提供更准确的更新,但偏向它们可能会低估其他客户端的数据。我们将这种权衡与实际多步局部更新的偏差和方差联系起来,产生一个采样优化问题和一个实用的维度感知近似,并直接进行模型平均。在语言任务上的实验考察了任务性能、客户端内存在数据异构性下的采样情况。结果表明,混合阶局部训练可以在显著降低客户端内存需求的情况下,保留大部分全FO性能。我们的代码可在该https URL获取。
英文摘要
Federated learning (FL) on memory-constrained edge devices faces a dilemma: first-order (FO) optimization (i.e., backpropagation) demands substantial memory, whereas zeroth-order (ZO) optimization suffers from severe convergence slowdown. To resolve this dilemma, we introduce HO-FL, a hybrid-order FL framework that trains a model's bottom segment with ZO optimization and its top segment with FO optimization. Each device can flexibly select its order boundary according to its memory budget while participating in the training of the same global model. Moreover, our convergence analysis reveals a new, fundamental trade-off: clients with larger FO-trained segments can provide more accurate updates, but favoring them can underrepresent other clients' data. We connect this trade-off to the bias and variance of actual multi-step local updates, yielding a sampling optimization problem and a practical dimension-aware approximation with direct model averaging. Experiments on language tasks examine task performance, client memory, and sampling under data heterogeneity. The results show that hybrid-order local training can retain much of the full-FO performance with substantially lower client memory requirements. Our code is available at https://github.com/HKU-WILL-Lab/HO-FL.
Comments30 pages, 2 figures