发表机构
Georgia Institute of Technology; George Mason University; Flatiron Institute(佐治亚理工学院; 乔治梅森大学; 弗拉蒂伦研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究指出联邦优化中的稳定性边缘(EoS)和渐进锐化是SCAFFOLD性能弱于FedAvg的原因,发现EoS会导致SCAFFOLD估计全局梯度的能力严重下降。
AI 中文摘要
在联邦学习中,众所周知异构数据(理论上)会减缓优化过程,大量研究致力于设计不受数据异构性影响的优化算法,如SCAFFOLD算法。然而尽管SCAFFOLD有强大的理论保证,在实际应用中它通常并不比简单得多的FedAvg算法表现更好。本研究通过大量实证探究,提出这种性能差距是由联邦优化中的稳定性边缘(Edge of Stability,EoS)和渐进锐化导致的。首先,我们发现在多种架构和超参数设置下,FedAvg和SCAFFOLD都会出现类似EoS的动态变化;我们观察到锐化的平衡值与学习率成反比(与梯度下降GD一致),且有趣的是,数据异构性程度(而非本地轮数)也会影响该平衡值。最重要的是,我们发现SCAFFOLD估计全局目标梯度的能力在EoS处严重下降,这通过优化轨迹上锐化程度与SCAFFOLD估计全局梯度误差的相关性来衡量。这为SCAFFOLD在深度学习中表现平平提供了一种机制:在EoS处高锐化的情况下,SCAFFOLD无法可靠地估计全局梯度。
英文摘要
In federated learning, it is well known that heterogeneous data can (in theory) slow down optimization, and much effort has been directed at designing optimization algorithms that are unaffected by data heterogeneity, such as the SCAFFOLD algorithm. Yet, despite strong theoretical guarantees, SCAFFOLD does not usually outperform the much simpler FedAvg in practice. In this work, we propose that this gap is due to the presence of Edge of Stability (EoS) and progressive sharpening in federated optimization, supported by extensive empirical probing. First, we find that EoS-like dynamics occur with both FedAvg and SCAFFOLD under a variety of architectures and hyperparameters. We observe that the equilibrium value of the sharpness is inversely proportional to the learning rate (as in GD), and interestingly, the degree of data heterogeneity (but not the number of local steps) also affects the equilibrium value. Most importantly, we observe that SCAFFOLD's ability to estimate the gradient of the global objective is severely degraded at the EoS, as measured by the correlation between sharpness and SCAFFOLD's error in estimating the global gradient along the optimization trajectory. This suggests a mechanism for SCAFFOLD's lackluster performance in deep learning: with high sharpness at the EoS, SCAFFOLD cannot reliably estimate the global gradient.
Comments7 pages, 7 figures