发表机构
University of Cambridge; University of Florida(剑桥大学; 佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出并行重启朗之万和朗之万-梯度混合方案,通过分离探索与利用,将全局优化的模拟时间从指数级降至多项式级,并在实验中验证了方法的有效性。
AI 中文摘要
我们研究了光滑、可能非凸的目标函数$\Gamma:\mathbb{R}^d\to\mathbb{R}$的全局优化所需的计算工作量。若算法的输出$\widehat X$满足$\mathbb{P}\{\Gamma(\widehat X)-\Gamma^\star>\varepsilon\}\leq\delta$,则该算法满足$(\varepsilon,\delta)$-PAC性能要求。算法设计与分析均在连续时间框架下进行。我们将经典模拟退火和固定温度朗之万扩散与本文引入并分析的两种方法进行比较:并行重启朗之万方法,以及使用随机动力学进行全局探索、梯度流进行局部利用的朗之万-梯度混合方案。设$L=\log(1/\delta)$,并令$E_*$表示主导能量势垒。在对数精度下,低温区域中前两种方法所需的模拟时间随$L/\varepsilon$呈指数增长。对于并行固定温度朗之万方法,当$\delta\downarrow0$且每个固定$\varepsilon>0$时,适当数量的独立试验给出$ C_3=L^{1+o(1)}/\varepsilon$。最显著的改进来自于将探索与利用分离。若$\eta$为包含全局最小化器的目标区域的吸引裕度,则在最优状态朗之万-梯度变体中,总模拟时间的充分低温估计为$ C_4^{(c)}\approx N\exp\{EL/(N\eta)\}+O(\log(1/\varepsilon))$,其中$E>E_*$。因此,全局探索与所要求的精度解耦。超越对数精度的分析揭示了维度相关的预因子,而对六峰驼峰和Rastrigin目标函数的实验则展示了更温暖探索的益处以及谱信息在理解探索时间方面的有用性。
英文摘要
This paper addresses a fundamental question in non-convex optimization: \textit{How should a stochastic optimizer allocate computation between global exploration and local exploitation?} We study this question in continuous time, using Langevin dynamics for global exploration and deterministic gradient flow for local exploitation. The simplest algorithm in this class is the best-state Langevin--gradient method: several Langevin trajectories explore the objective landscape, the best state encountered is retained, and a single gradient trajectory then refines this state to high terminal accuracy. Our main objective is to characterize the computational work required to achieve a prescribed accuracy with prescribed confidence. We develop low-temperature approximations for this probably approximately correct (PAC) work--accuracy tradeoff, both for the best-state method and for related Langevin schemes. Global exploration is governed by energy barriers, through spectral-gap and Eyring--Kramers asymptotics. The resulting approximations expose the competing effects of temperature, computational budget, and the number of exploratory trajectories, and quantify their dependence on the confidence parameter. In contrast, terminal accuracy is largely separated from global exploration: once a suitable region of attraction has been reached, gradient flow requires only $O(\log(1/\varepsilon))$ additional work to reach accuracy at least $\varepsilon$. Numerical experiments on challenging objectives serve to compare algorithms and to test the quantitative predictions of the asymptotic theory. They identify regimes in which the low-temperature approximations provide useful guidance, as well as regimes exhibiting substantial temperature sensitivity and potentially significant high-dimensional limitations.