发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究非凸-强凸随机双层优化中寻找超梯度范数小于ε的点的样本复杂度,证明了在固定正则性预算下极小极大复杂度为Θ_p(ε^{-4-2/p}),并刻画了高阶光滑性下的精确复杂度,表明定量光滑性带来精确收益。
AI 中文摘要
我们研究了在非凸-强凸随机双层优化中,使用有界方差的全新全局无偏一阶样本,寻找期望超梯度范数至多 $\epsilon>0$ 的点的样本复杂度。对于下变量光滑阶 $p\ge1$,F$^2$SA-$p$ 实现了 $\widetilde O(p\epsilon^{-4-2/p})$(Chen 等人,2026),但指数 $4+2/p$ 的必要性此前是开放的。我们在固定的非退化正则性预算下证明了匹配的上界和下界 $\Theta_p(\epsilon^{-4-2/p})$,其中常数允许依赖于 $p$。下界在固定样本预算下对无限制随机算法成立,允许维度随精度增长,并保持全局无偏性和所有规定的下变量光滑界,即使在下变量为标量且上梯度精确的情况下也成立。我们进一步刻画了在 $L_j=M\Lambda^j(j!)^\beta$($j\ge3$)下的极小极大固定预算复杂度:$Q_{p,\beta}^*(\epsilon)=\Theta\left(\epsilon^{-4}\left[\inf_{1\le r\le p,,r\in\mathbb N} r^\beta\epsilon^{-1/r}\right]^2\right)$,其中 $p\in\mathbb N\cup\{\infty\}$,常数与 $p$ 无关。在无限阶时,这给出 $\beta=0$ 时的 $\Theta(\epsilon^{-4})$ 和 $\beta=1$ 时的 $\Theta(\epsilon^{-4}\log^2(1/\epsilon))$。一种使用加权独立批次的梯度差分方法在阶数上一致地达到这些界。因此,定量光滑性带来精确的收益,而仅定性解析性仍允许 $\Theta(\epsilon^{-6})$ 的复杂度。
英文摘要
We study the sample complexity of finding a point with expected hypergradient norm at most $ε>0$ in nonconvex--strongly-convex stochastic bilevel optimization using fresh, globally unbiased first-order samples with bounded variance. For lower-variable smoothness order $p\ge1$, F$^2$SA-$p$ achieves $\widetilde O(pε^{-4-2/p})$ (Chen et al., 2026), but necessity of the exponent $4+2/p$ was open. We prove matching upper and lower bounds $Θ_p(ε^{-4-2/p})$ under fixed nondegenerate regularity budgets, with constants allowed to depend on $p$. The lower bound holds for unrestricted randomized algorithms under a fixed sample budget, allows dimension to grow with accuracy, and preserves global unbiasedness and all prescribed lower-variable smoothness bounds, even with a scalar lower variable and exact upper gradients. We further characterize the minimax fixed-budget complexity under $L_j=MΛ^j(j!)^β$ for $j\ge3$: $Q_{p,β}^*(ε)=Θ\left(ε^{-4}\left[\inf_{1\le r\le p,,r\in\mathbb N} r^βε^{-1/r}\right]^2\right)$, for $p\in\mathbb N\cup{\infty}$, with constants independent of $p$. At infinite order, this gives $Θ(ε^{-4})$ for $β=0$ and $Θ(ε^{-4}\log^2(1/ε))$ for $β=1$. A gradient-difference method with weighted independent batches attains these bounds uniformly in order. Thus quantitative smoothness yields precise gains, whereas qualitative analyticity alone still permits $Θ(ε^{-6})$ complexity.