发表机构
MIRIAI; HSE University; MBZUAI; Innopolis University; AI Institute MSU(MIRIAI; 高等经济大学; 穆罕默德·本·扎耶德人工智能大学; 喀山国立研究技术大学; 莫斯科国立大学人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究高度光滑凸函数的零阶优化,提出加速有限差分方法,在强凸情形下达到调用次数下界 $d^2\varepsilon^{-\beta/(\beta-1)}$ 且轮次为 $O(\sqrt{\kappa}\log(1/\varepsilon))$,并证明可容忍最优噪声,实验验证理论。
AI 中文摘要
我们研究了凸函数和强凸函数的零阶优化问题,这些函数具有阶数 $\beta>2$ 的 Hölder 光滑性,其函数值受到随机噪声的污染,该随机噪声的偏差与方法的随机方向无关,同时存在 Lipschitz 系统性误差和有界对抗性误差。我们衡量了函数评估(调用)的次数、并行调用的顺序轮次,以及可容忍的最大系统性和对抗性误差。对于具有有界 Hessian 的强凸函数,我们证明了调用次数的下界为 $d^2\varepsilon^{-\beta/(\beta-1)}$ 阶,其中 $d$ 是维度,$\varepsilon$ 是精度,并在更强的 Hilbert--Schmidt 光滑性假设下,对所有 $\beta>2$ 找到了 $d$ 的精确幂次。一种加速有限差分方法在条件数 $\kappa$ 下,以 $O(\sqrt{\kappa}\log(1/\varepsilon))$ 轮次达到该下界,这与使用精确梯度的 Nesterov 方法相同,并且能容忍最优的系统性误差以及(在对数因子内)最优的对抗性误差;速率最优的方法需要 $\Omega(\log d)$ 轮次。当 Hessian 随梯度至多线性增长时,以及对于中心化噪声且 Hessian 无上界的情况,也能在对数轮次内达到相同的调用次数。对于凸函数,我们的下界在 $d$ 上多项式地改进了已知结果,并且在困难实例的曲率处是紧的;系统性误差水平同样是最优的。这些界推广到一般范数和约束;在 $\ell_1$ 球上,镜像下降所需的调用次数在 $d$ 上接近线性。实验证实了该理论。
英文摘要
We study zeroth-order optimization of convex and strongly convex functions with Hölder smoothness of order $β>2$ from values corrupted by random noise with bias independent of the method's random directions, a Lipschitz systematic error and a bounded adversarial error. We measure the number of function evaluations (calls), of sequential rounds of parallel calls, and the largest tolerable systematic and adversarial errors. For strongly convex functions with a bounded Hessian, we prove a lower bound of order $d^2\varepsilon^{-β/(β-1)}$ on the number of calls, where $d$ is the dimension and $\varepsilon$ the accuracy, and find the sharp power of $d$ under the stronger Hilbert--Schmidt smoothness for all $β>2$. An accelerated finite-difference method attains them in $O(\sqrtκ\log(1/\varepsilon))$ rounds for condition number $κ$, like Nesterov's method with exact gradients, and tolerates the optimal systematic and, up to logarithms, adversarial errors; rate-optimal methods need $Ω(\log d)$ rounds. The same number of calls is attained when the Hessian grows at most linearly with the gradient and, for centered noise, without an upper bound on the Hessian, also in logarithmically many rounds. For convex functions, our lower bound improves the known one polynomially in $d$ and is tight at the curvature of the hard instances; the systematic level is again optimal. These bounds extend to general norms and constraints; on the $\ell_1$ ball, mirror descent needs calls nearly linear in $d$. Experiments confirm the theory.