超越传统联邦学习:基于高阶正则化
Beyond Conventional Federated Learning via High-Order Regularization
- University of Antwerp(安特卫普大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出HiFedProx,用幂型正则化替代二次惩罚,通过调节指数p压缩客户端位移差异,在FEMNIST上显著提升压力场景性能,建议校准指数而非最大化。
AI中文摘要:
联邦客户端执行多次本地优化步骤后,返回的参数位移幅度可能差异巨大。FedProx的二次正则化随位移线性增长,因此对普通客户端与异常大客户端移动之间的对比控制有限。本文提出HiFedProx,将二次惩罚替换为以$p\geq2$为索引的尺度匹配幂型正则化器。所有幂次在参考位移$R$处具有相同的正则化梯度幅度,而每个$p>2$在低于$R$时响应较弱,高于$R$时响应较强。精确的仿射参考计算表明,增大$p$会压缩相对位移差异,尽管非常大的幂次接近固定半径行为并增加局部曲率。HiFedProx将该几何特性与有限预算随机客户端优化及同小批量Armijo回溯相结合。在冻结的60位书写者FEMNIST子集上进行的配对五种子实验中,对$p\in\{2,3,4,5,6,7,8\}$的公共参数研究表明,在干净训练下性能相似,但在复合压力下性能显著提升。中等和严重压力下的最低损失分别出现在$p=7$和$p=6$,相比$p=2$分别提高了$11.44\\%$和$23.16\\%$。尽管位移尾部比率在$p=8$时继续下降,但预测性能在中间范围达到峰值,且Armijo试验成本随$p$增加。这些结果表明,指数应被校准而非最大化。在我们的实验中,$p=5$--$7$提供了最有用的范围。
英文摘要:
Federated clients that perform several local optimization steps can return parameter displacements with widely different magnitudes. The quadratic regularization of FedProx grows linearly with displacement and therefore offers limited control over the contrast between ordinary and unusually large client movements. We here introduce HiFedProx, which replaces the quadratic penalty with a scale-matched power-type regularizer indexed by $p\geq2$. All powers have the same regularization-gradient magnitude at a reference displacement $R$, while every $p>2$ gives a weaker response below $R$ and a stronger response above it. An exact affine reference calculation shows that increasing $p$ compresses relative displacement disparities, although very large powers approach fixed-radius behavior and increase local curvature. HiFedProx combines this geometry with finite-budget stochastic client optimization and same-minibatch Armijo backtracking. In paired five-seed experiments on a frozen 60-writer FEMNIST subset, a common-parameter study over $p\in\{2,3,4,5,6,7,8\}$ shows similar clean-training performance but substantial gains under composite stress. The lowest moderate- and severe-stress losses occur at $p=7$ and $p=6$, improving over $p=2$ by $11.44\%$ and $23.16\%$, respectively. Although displacement-tail ratios continue to decrease through $p=8$, predictive performance peaks in an intermediate range and Armijo trial cost increases with $p$. These results indicate that the exponent should be calibrated rather than maximized. In our experiments, $p=5$--$7$ provides the most useful range.