发表机构
University of Oxford; University of Economics and Business Vienna; University of Groningen(牛津大学; 维也纳经济与商业大学; 格罗宁根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种仅对DBGD乘子进行分母正则化的方法,在无需罕见访问假设下实现随机非凸双层优化的快速收敛,达到(ε, ε)-平稳性,改进了无假设复杂度。
AI 中文摘要
我们研究了具有光滑且可能非凸的上层和下层目标的随机简单双层优化。现有的动态障碍梯度下降(DBGD)的随机扩展要么在不可验证的、依赖于轨迹的“罕见访问”假设下获得快速收敛,要么以显著更高的预言机成本去除该假设。我们证明了仅对DBGD乘子进行分母正则化即可消除此类假设的需求,同时保持快速收敛速率。具体而言,我们的方法在O(ε^{-2})次迭代中达到(ε, ε)-平稳性,使用O(ε^{-4})个上层和O(ε^{-7})个下层随机梯度,这改进了最佳的无假设复杂度。我们还推导了任意时间参数调度。
英文摘要
We investigate stochastic simple bilevel optimization with smooth and possibly nonconvex upper- and lower-level objectives. Existing stochastic extensions of dynamic barrier gradient descent (DBGD) either obtain fast convergence under an unverifiable trajectory-dependent ``rare-visit'' assumption, or remove this assumption at a substantially higher oracle cost. We show that a simple denominator-only regularization of the DBGD multiplier eliminates the need for such an assumption while preserving fast convergence rates. Specifically, our method achieves $(\varepsilon, \varepsilon)$-stationarity in $O(\varepsilon^{-2})$ iterations using $O(\varepsilon^{-4})$ upper-level and $O(\varepsilon^{-7})$ lower-level stochastic gradients, which improves upon the best assumption-free complexities. We additionally derive anytime parameter schedules.