arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16350cs.LG

联邦随机双层优化:仅需一阶梯度的方法

Federated stochastic bilevel optimization with fully first-order gradients

  • Temple University(天普大学)
  • Stony Brook University(石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

Yihan Zhang, Rohit Dhaipule, Chiu C Tan, Haibin Ling, Hongchang Gao

AI总结:

针对联邦随机双层优化中计算二阶矩阵导致运行时间长的问题,提出仅用一阶梯度并采用常数单时间尺度学习率的方差缩减算法,显著降低运行时间,实验验证有效。

AI中文摘要:

近年来,由于在机器学习中的广泛应用,联邦随机双层优化受到了积极研究。然而,大多数现有的联邦随机双层优化算法需要计算二阶Hessian和Jacobian矩阵,这导致实际运行时间较长。为应对这些挑战,我们提出了一种新颖的联邦随机方差缩减双层梯度下降算法,该算法仅依赖一阶预言机。具体而言,我们的方法不需要计算二阶Hessian和Jacobian矩阵,从而显著减少了运行时间。此外,我们引入了一种新颖的学习率机制,即常数单时间尺度学习率,以协调不同变量的更新。我们还提出了一种新策略来建立我们算法的收敛率。最后,广泛的实验结果证实了我们所提出算法的有效性。

英文摘要:

Federated stochastic bilevel optimization has been actively studied in recent years due to its widespread applications in machine learning. However, most existing federated stochastic bilevel optimization algorithms require the computation of second-order Hessian and Jacobian matrices, which leads to longer running times in practice. To address these challenges, we propose a novel federated stochastic variance-reduced bilevel gradient descent algorithm that relies solely on first-order oracles. Specifically, our approach does not require the computation of second-order Hessian and Jacobian matrices, significantly reducing running time. Furthermore, we introduce a novel learning rate mechanism, i.e., a constant single-timescale learning rate, to coordinate the update of different variables. We also present a new strategy to establish the convergence rate of our algorithm. Finally, the extensive experimental results confirm the efficacy of our proposed algorithm.

补充信息

↑