arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高度光滑强凸零阶优化的尖锐极小极大速率

Sharp Minimax Rates for Highly Smooth Strongly Convex Zeroth-Order Optimization

Haihan Zhang, Wendao Wu, Chenheng Zhang, Yanyi Li, Chunyuan Zheng, Cong Fang, Haoxuan Li, Zhouchen Lin

arXiv 2610.00276首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究确定了高阶光滑强凸函数在带噪零阶优化中的极小极大维度依赖速率,并给出了匹配的上下界,表明查询复杂度为 $\Theta(d^2\varepsilon^{-\beta/(\beta-1)})$。

AI 中文摘要

我们确定了在多重线性算子范数下具有高阶光滑性的全局强凸函数的带噪零阶优化的极小极大维度依赖关系。每次查询返回一个带有独立高斯噪声的函数值。对于维度 $d$、查询预算 $T$ 以及任意固定的有限光滑阶数 $\beta>2$,极小极大期望目标误差满足 $\mathcal{E}_{\beta}(T,d)=\Theta_{\beta}\left(\min\left\{1,(d^2/T)^{(\beta-1)/\beta}\right\}\right)$,其中强凸性、梯度光滑性、高阶光滑性、噪声以及极小点半径的界均与维度无关。因此,对于足够小的目标误差 $\varepsilon$,必要且充分的带噪查询次数为 $\Theta_{\beta}(d^2\varepsilon^{-\beta/(\beta-1)})$,且无对数损失。新的下界与 Akhavan 等人(2024)的上界在维度和预算依赖上相匹配,并适用于具有无限制查询点的任意随机自适应算法。其关键步骤是使用一个径向截断来局部化归一化超立方体:一次查询可以强调一个隐藏坐标,但关于所有坐标的聚合信息仍然很小。一个自包含的多尺度球面估计器通过抵消低阶偏差项达到匹配的上界。速率中的常数仅依赖于所述归一化下的固定光滑阶数。

英文摘要

We determine the minimax dimension dependence of noisy zeroth-order optimization for globally strongly convex functions with higher-order smoothness in multilinear operator norm. Each query returns one function value with fresh independent Gaussian noise. For dimension $d$, query budget $T$, and any fixed finite smoothness order $β>2$, the minimax expected objective error satisfies $\mathcal{E}_β(T,d)=Θ_β\left(\min\left\{1,(d^2/T)^{(β-1)/β}\right\}\right)$, with dimension-independent strong-convexity, gradient-smoothness, higher-order-smoothness, noise, and minimizer-radius bounds. Thus, for sufficiently small target error $\varepsilon$, the necessary and sufficient number of noisy values is $Θ_β(d^2\varepsilon^{-β/(β-1)})$, with no logarithmic loss. The new lower bound matches the dimension and budget dependence of the upper bound of Akhavan et al. (2024), and holds for arbitrary randomized adaptive algorithms with unrestricted query points. Its key step localizes a normalized hypercube with one radial cutoff: a query can emphasize one hidden coordinate, but the aggregate information about all coordinates remains small. A self-contained multiscale spherical estimator attains the matching upper bound by cancelling lower-order bias terms. Constants in the rates depend only on the fixed smoothness order under the stated normalization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑