缩小随机一阶双层优化的精度差距
Closing the Accuracy Gap in Stochastic First-Order Bilevel Optimization
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对随机一阶双层优化,证明非凸-强凸情形下ε^{-6}的样本复杂度下界,并给出匹配的冻结线性倾斜估计器,缩小了精度差距。
AI中文摘要:
我们刻画了使用随机一阶预言机的光滑非凸-强凸双层优化的最优精度依赖关系。对于具有有界方差的全局无偏新鲜梯度观测,我们证明了与现有上界精度指数相匹配的Ω(ε^{-6})下界。我们考虑全局Lipschitz梯度、下变量中的有界上梯度以及Lipschitz下Hessian块,目标为E|∇F(x̂)|≤ε。设L为共同一阶尺度,μ为下层强凸模量,ρ为下层Hessian变化预算,κ=L/μ,χ=1+ρ/μ。对于ρ≳μ、χ≲κ、足够高的精度和足够大的维度,我们建立了Ω(Lκ²Δε^{-2}+L³χ²κ⁸σ_g²Δε^{-6}),其中Δ为初始差距预算,σ_g²为下层梯度方差预算。该界在完整欧几里得空间上对任意随机自适应算法成立,即使每个样本都是标量函数的梯度。在共同尺度区间ρ=Θ(L)中,主要随机项的条件数依赖为κ^{10},而可实现依赖为κ^{11}。证明将一个序列硬目标嵌入到精确的下层响应中,同时保持全局正则性和有限差距。多尺度分解限制了携带每个新方向的梯度信号,从而得出所需的样本复杂度。我们用一个冻结线性倾斜估计器补充该下界,其偏差与下层Hessian变化成正比。其分析使这一结构依赖明确化,并在指定的曲率控制和小差距区间中得出匹配的主要速率。
英文摘要:
We characterize the optimal accuracy dependence for smooth nonconvex--strongly-convex bilevel optimization with stochastic first-order oracles. For globally unbiased fresh-gradient observations with bounded variance, we prove an $Ω(ε^{-6})$ lower bound matching the accuracy exponent of existing upper bounds. We consider globally Lipschitz gradients, a bounded upper gradient in the lower variable, and Lipschitz lower Hessian blocks, with target $\mathbb{E}|\nabla F(\widehat{x})|\leε$. Let $L$ be the common first-order scale, $μ$ the lower strong-convexity modulus, $ρ$ the lower Hessian-variation budget, $κ=L/μ$, and $χ=1+ρ/μ$. For $ρ\gtrsimμ$, $χ\lesssimκ$, sufficiently high accuracy, and sufficiently large dimension, we establish $Ω\left(Lκ^2Δε^{-2}+L^3χ^2κ^8σ_g^2Δε^{-6}\right)$, where $Δ$ is the initial gap budget and $σ_g^2$ is the lower-gradient variance budget. The bound holds over full Euclidean spaces against arbitrary randomized adaptive algorithms, even when every sample is the gradient of a scalar function. In the common-scale regime $ρ=Θ(L)$, the leading stochastic term has condition-number dependence $κ^{10}$, compared with the achievable $κ^{11}$ dependence. The proof embeds a sequential hard objective into an exact lower response while preserving global regularity and finite gap. A multiscale decomposition limits the gradient signal carrying each new direction, yielding the required sample complexity. We complement this lower bound with a frozen-linear-tilt estimator whose bias is proportional to lower Hessian variation. Its analysis makes this structural dependence explicit and yields matching leading rates in the specified curvature-controlled and small-gap regimes.