发表机构
Texas A&M University(得克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对随机双层优化,提出单循环恒定批量一阶罚方法SICO,结合投影与指数移动平均,实现最优样本复杂度,解决开放问题。
AI 中文摘要
随机双层优化(SBO)的罚方法近期进展已消除对二阶导数预言机的需求。然而,对于随机非凸-强凸双层问题,现有的一阶方法通常依赖嵌套循环和/或大批量,以在标准有界方差假设或均方光滑性假设下获得 $O(\epsilon^{-6})$ 或 $O(\epsilon^{-4})$ 的样本复杂度。由于精确逼近需要较大的罚值,使用单循环罚方法和恒定批量实现这些速率仍然具有挑战性。为解决这一挑战,我们开发了一种随机单循环恒定批量一阶罚方法(SICO),它结合了两个互补的要素。首先,它对原始下层问题和罚问题每次迭代执行一次随机梯度更新,并通过投影控制其迭代之间的分离。其次,它应用指数移动平均来稳定上层梯度估计器。我们证明,在无偏、有界方差的随机梯度下,该组合仅使用每次迭代 $O(1)$ 个随机梯度样本即可实现 $O(\epsilon^{-6})$ 的样本复杂度。在下层随机梯度的额外均方光滑性假设下,同一算法将复杂度提升至 $O(\epsilon^{-4})$,且批量大小仍为 $O(1)$。据我们所知,这是首个使用单循环和恒定批量匹配完全一阶SBO方法最知名收敛速率的工作。该结果解决了文献中提出的一个开放问题。
英文摘要
Recent advances in penalty-based methods for stochastic bilevel optimization (SBO) have eliminated the need for second-order derivative oracles. However, for stochastic nonconvex-strongly convex bilevel problems, existing first-order methods typically rely on nested loops and/or large batch sizes for attaining $O(ε^{-6})$ or $O(ε^{-4})$ sample complexity under standard bounded-variance assumption or mean-square smoothness assumption. Achieving these rates with a single-loop penalty method and a constant batch size remains challenging due to a large penalty value needed for an accurate approximation. To address this challenge, we develop a stochastic SIngle-loop COnstant-Batch first-order penalty method (SICO) that combines two complementary ingredients. First, it performs one stochastic-gradient update per-iteration for both the original lower-level and penalized problems, with a projection that controls the separation between their iterates. Second, it applies an exponential moving average to stabilize the upper-level gradient estimator. We show that this combination achieves $ O(ε^{-6}) $ sample complexity using only $O(1)$ stochastic-gradient samples per iteration under unbiased, bounded-variance stochastic gradients. Under the additional mean-square smoothness assumption on the lower-level stochastic gradients, the same algorithm improves the complexity to $O(ε^{-4})$ also with $O(1)$ batch size. To the best of our knowledge, this is the first work to match the best-known convergence rate for fully first-order SBO methods using a single loop and a constant batch size. This result addresses an open problem posed in the literature.