发表机构
École polytechnique; Courant Institute, NYU(巴黎综合理工学院; 纽约大学柯朗研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明当梯度近似成本随精度超线性增长时,用随机化多水平预言机驱动不精确梯度下降,可使凸优化成本降至与单次梯度评估同阶,并显著提升收敛速度。
AI 中文摘要
本文研究了当梯度无法精确评估,而只能通过一系列算法来近似时,这些算法的计算量随精度δ的增长而像δ^{-γ}那样增长。当γ>2时,进入“比蒙特卡洛更难”(HTMC)区域,精度的代价超过了蒙特卡洛所能带来的方差缩减,我们证明最小化损失函数的成本,最多不超过在问题所需精度下对梯度进行一次评估的成本,其倍数仅取决于γ。一个随机化的多水平预言机用其无偏估计器取代了精度δ的确定性近似,该估计器的方差σ²成为第二个独立定价的旋钮:一次调用的成本从δ^{-γ}降至δ^{2-γ}σ^{-2}。由该预言机驱动的普通不精确梯度下降,在凸情形下达到损失ε的期望计算量为Θ(ε^{-γ}),而在固定精度下运行相同方法则需要Θ(ε^{-(γ+1)}):随机化买来了ε的一个完整幂次。在μ-强凸条件下,指数减半,变为ε^{-γ/2},因为迭代序列稳定在噪声底板上,偏差预算相应放宽。两个界都与步长无关,因此也与光滑性常数无关,我们证明成本是底层梯度流的泛函,而非其任何离散化的泛函。
英文摘要
This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $δ^{-γ}$ in the accuracy $δ$. When $γ>2$, falling into the Harder-Than-Monte-Carlo (HTMC) regime, the price of accuracy outruns the variance reduction that Monte Carlo would buy and we show that minimizing a loss function costs no more, up to a factor depending only on $γ$, than a single evaluation of its gradient at the accuracy the problem demands. A randomized multilevel oracle replaces the deterministic approximation of accuracy $δ$ by an unbiased estimator of it, whose variance $σ^2$ becomes a second, independently priced dial: the cost of one call drops from $δ^{-γ}$ to $δ^{2-γ}σ^{-2}$. Plain inexact gradient descent driven by that oracle reaches loss $\varepsilon$ at expected compute $Θ(\varepsilon^{-γ})$ in the convex case, against $Θ(\varepsilon^{-(γ+1)})$ for the same method run at a fixed accuracy: randomization buys a full power of $\varepsilon$. Under $μ$-strong convexity the exponent halves, to $\varepsilon^{-γ/2}$, because the iterates settle at a noise floor and the bias budget relaxes accordingly. Both bounds are independent of the step size, and hence of the smoothness constant, and we show that the cost is a functional of the underlying gradient flow rather than of any discretization of it.