arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全局收敛的三阶朗之万动力学用于非凸优化:基于模拟退火

Global Convergence of Third-Order Langevin Dynamics for Non-Convex Optimization via Simulated Annealing

Yingli Wang, Lingjiong Zhu

arXiv 2609.28611首次发表:更新:

发表机构

Fudan University; Florida State University(复旦大学; 佛罗里达州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出三阶朗之万动力学结合模拟退火实现非凸优化全局收敛,通过扭曲熵和三次端点估计保证速率,数值实验显示优于UBU和过阻尼方法。

AI 中文摘要

我们研究了通过模拟退火、固定摩擦和递减噪声进行非凸优化的三阶朗之万动力学的全局收敛保证。一个显式的三块扭曲熵将耗散从噪声辅助变量转移到完整状态。在耗散性、正则性和低温函数不等式假设下,对数冷却以势垒控制的动力学速率驱动目标值在概率上达到全局最小值。对于精确力积分和中点三阶段离散化,多项式递减步长在物理时间尺度上保持该速率。三次局部端点估计给出了比现有冻结力动力学结果更宽松的充分步长条件。与单梯度UBU积分器的比较表明,在相同的强耦合分析下,其中心化随机局部误差导致更小的充分迭代指数。进行了数值实验以说明我们的理论。对于双阱目标,三阶朗之万终端成功点估计在共同视界和相等梯度预算下均高于UBU。对于使用合成数据的高维非凸神经网络目标,独立调优的UBU和三阶朗之万方案均优于过阻尼朗之万动力学;三阶朗之万点估计更高。对于相同神经网络目标在真实数据上,我们展示了最佳盆地概率和淬火后测试准确率的相同点估计排序。数值代码和相关实验结果可在https URL公开获取。

英文摘要

We study global convergence guarantees of third-order Langevin dynamics for non-convex optimization via simulated annealing with fixed friction and decreasing noise. An explicit three-block distorted entropy transfers dissipation from the noisy auxiliary variable to the full state. Under dissipativity, regularity, and low-temperature functional-inequality assumptions, logarithmic cooling drives the objective values to the global minimum in probability at the barrier-controlled kinetic rate. For the exact-force-integral and midpoint three-stage discretizations, polynomially decreasing steps preserve this rate on the physical time scale. The cubic local endpoint estimate gives a less restrictive sufficient step-size condition than the available frozen-force kinetic result. A comparison with the one-gradient UBU integrator shows how its centered stochastic local error leads, under the same strong-coupling analysis, to a smaller sufficient iteration exponent. Numerical experiments are conducted to illustrate our theory. For a double well objective, third-order Langevin terminal-success point estimates are higher than UBU at both a common horizon and an equal gradient budget. For a high-dimensional nonconvex neural-network objective using synthetic data, independently tuned UBU and third-order Langevin schemes both outperform overdamped Langevin dynamics; the third-order Langevin point estimate is higher. For the same neural-network objective on real data, we show the same point-estimate ordering for best-basin probability and post-quench test accuracy. Numerical code and associated experiment results are publicly available at https://github.com/gagawjbytw/simulated-annealing-third-order-langevin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑