arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于非凸优化的朗之万方法:精确、不精确和零阶

Langevin for Nonconvex Optimization: Exact, Inexact and Zeroth-Order

Emanuele Naldi, Marco Rando, Lorenzo Rosasco, Silvia Villa

arXiv 2607.22353首次发表:更新:

发表机构

University of Genova; Université Côte d’Azur; INRIA; CNRS; Istituto Italiano di Tecnologia(热那亚大学; 蔚蓝海岸大学; 法国国家数字研究院; 法国国家科学研究中心; 意大利理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于朗之万的非凸优化方法,通过特定不等式和估计从相对熵到目标值误差,先分析精确梯度的未调整朗之万算法,再扩展到不精确梯度版本,给出零阶朗之万优化复杂度界并进行数值实验。

AI 中文摘要

我们研究在平滑性和耗散性假设下基于朗之万的非凸优化方法。重点是获得预期超额风险的非渐近界而非整个目标分布的采样保证。分析关键是基于加权的Csiszár - Kullback - Pinsker不等式和指数矩估计,直接从相对熵过渡到目标值误差,避免中间的瓦瑟斯坦界。首先分析带精确梯度的未调整朗之万算法,得出关于\(\mathbb{E}[F(x_k)] - \min F\)的显式界,然后扩展到不精确梯度版本,涵盖随机梯度和仅基于函数评估的零阶估计器,给出零阶朗之万优化的显式函数评估复杂度界,还进行了数值实验说明所提零阶朗之万方案的行为。

英文摘要

We study Langevin-based methods for non-convex optimization under smoothness and dissipativity assumptions, focusing on non-asymptotic expected excess-risk bounds rather than sampling guarantees for the full target distribution. A main methodological message is that relative-entropy sampling guarantees can be converted directly into expected objective-value guarantees, without passing through Wasserstein distance. This direct KL-to-objective route yields sharper dependence on the Log-Sobolev constant, which may scale exponentially with inverse temperature and dimension in non-convex problems. We first analyze the Unadjusted Langevin Algorithm with exact gradients and derive explicit bounds in terms of inverse temperature, dimension, stepsize, smoothness and dissipativity parameters, and the Log-Sobolev constant. We then extend the result to inexact ULA, allowing biased and stochastic gradient surrogates whose mean-square error grows at most quadratically in the state. This framework covers stochastic gradients and zeroth-order estimators based on function evaluations. We show that Gaussian and spherical finite-difference estimators fit the theory and obtain explicit function-evaluation complexity bounds. To the best of our knowledge, these are the first non-asymptotic global non-convex optimization complexity bounds for zeroth-order ULA. We also provide numerical experiments illustrating the spherical zeroth-order scheme.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑