发表机构
University of Genova; Université Côte d’Azur; INRIA; CNRS; Istituto Italiano di Tecnologia(热那亚大学; 蔚蓝海岸大学; 法国国家数字研究院; 法国国家科学研究中心; 意大利理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于朗之万的非凸优化方法,通过特定不等式和估计从相对熵到目标值误差,先分析精确梯度的未调整朗之万算法,再扩展到不精确梯度版本,给出零阶朗之万优化复杂度界并进行数值实验。
AI 中文摘要
我们研究在平滑性和耗散性假设下基于朗之万的非凸优化方法。重点是获得预期超额风险的非渐近界而非整个目标分布的采样保证。分析关键是基于加权的Csiszár - Kullback - Pinsker不等式和指数矩估计,直接从相对熵过渡到目标值误差,避免中间的瓦瑟斯坦界。首先分析带精确梯度的未调整朗之万算法,得出关于\(\mathbb{E}[F(x_k)] - \min F\)的显式界,然后扩展到不精确梯度版本,涵盖随机梯度和仅基于函数评估的零阶估计器,给出零阶朗之万优化的显式函数评估复杂度界,还进行了数值实验说明所提零阶朗之万方案的行为。
英文摘要
We study Langevin-based methods for non-convex optimization under smoothness and dissipativity assumptions, focusing on non-asymptotic expected excess-risk bounds rather than sampling guarantees for the full target distribution. A main methodological message is that relative-entropy sampling guarantees can be converted directly into expected objective-value guarantees, without passing through Wasserstein distance. This direct KL-to-objective route yields sharper dependence on the Log-Sobolev constant, which may scale exponentially with inverse temperature and dimension in non-convex problems. We first analyze the Unadjusted Langevin Algorithm with exact gradients and derive explicit bounds in terms of inverse temperature, dimension, stepsize, smoothness and dissipativity parameters, and the Log-Sobolev constant. We then extend the result to inexact ULA, allowing biased and stochastic gradient surrogates whose mean-square error grows at most quadratically in the state. This framework covers stochastic gradients and zeroth-order estimators based on function evaluations. We show that Gaussian and spherical finite-difference estimators fit the theory and obtain explicit function-evaluation complexity bounds. To the best of our knowledge, these are the first non-asymptotic global non-convex optimization complexity bounds for zeroth-order ULA. We also provide numerical experiments illustrating the spherical zeroth-order scheme.