arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非凸联邦随机双层优化的平滑梯度方法

Smoothed Gradient Method for Nonconvex Federated Stochastic Bilevel Optimization

Xinwen Zhang, Peiran Yu, Zhaosong Lu, Hongchang Gao

arXiv 2610.05290首次发表:更新:

发表机构

Temple University; University of Texas at Arlington; University of Minnesota(天普大学; 德克萨斯大学阿灵顿分校; 明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非凸联邦随机双层优化,提出一种随机双重平滑梯度方法,解耦上下层学习率,无需强凸下层,收敛率与通信复杂度均优于现有方法。

AI 中文摘要

近年来,联邦随机双层优化因其在机器学习中的广泛应用而受到越来越多的关注。为减少与二阶Hessian和Jacobian矩阵相关的计算开销,已有多种一阶方法被提出。然而,现有方法通常对下层函数施加限制性假设,其收敛速度对条件数有较强依赖,并且需要对上层和下层问题的变量使用不同的学习率尺度,这限制了它们的实际适用性并使超参数调优复杂化。为应对这些挑战,我们提出了一种用于非凸联邦随机双层优化问题的随机双重平滑梯度方法,该方法解耦了上层和下层变量的学习率,且不需要下层损失函数为强凸。我们为所提算法建立了严格的理论保证,展示了改进的收敛率$O(\kappa^{15/2}/\epsilon^5)$和通信复杂度$O(\kappa^{4}/\epsilon^3)$,其中$\kappa$表示条件数,$\epsilon$表示解精度。值得注意的是,这些界对条件数$\kappa$的依赖显著优于现有方法。大量实验验证了我们算法的有效性。

英文摘要

In recent years, federated stochastic bilevel optimization has attracted increasing attention due to its wide range of applications in machine learning. To reduce the computational overhead associated with second-order Hessian and Jacobian matrices, several first-order methods have been proposed. However, existing methods typically impose restrictive assumptions on the lower-level function, suffer from a strong dependence on the condition number in their convergence rates, and require different learning-rate scales for variables across the upper- and lower-level problems, limiting their practical applicability and complicating hyperparameter tuning. To address these challenges, we propose a stochastic doubly smoothed gradient method for nonconvex federated stochastic bilevel optimization problems, which decouples the learning rates of upper- and lower-level variables and does not require a strongly-convex lower-level loss function. We establish rigorous theoretical guarantees for the proposed algorithm, demonstrating an improved convergence rate of $O(κ^{15/2}/ε^5)$ and a communication complexity of $O(κ^{4}/ε^3)$, where $κ$ denotes the condition number and $ε$ represents the solution accuracy. Notably, these bounds exhibit significantly better dependence on the condition number $κ$ than those of existing methods. Extensive experiments validate the effectiveness of our algorithm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑