发表机构
Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对梯度难处理的损失函数,提出一类基于扩散的随机粒子优化方法,引入平均场动力学及其交互粒子近似,证明其指数收敛性与非渐近误差界,并通过变体在相关任务上验证了有效性。
AI 中文摘要
我们针对梯度难以处理的损失函数,开发了一类基于扩散的随机粒子优化方法。具体而言,我们考虑损失梯度是关于参数依赖分布的积分的问题,这类问题涵盖生成模型训练、微调以及隐变量模型学习等场景。我们引入了平均场动力学及其交互粒子近似,该近似将多种现有算法作为特例,同时为构建新方法提供了途径。在适定性和联合收缩性假设下,我们证明了指数收敛性,并表明连续时间粒子系统具有非渐近误差界。我们通过开发动量和高阶朗之万变体,并在最大边际似然估计及基于能量的模型训练上对其进行评估,验证了该方法的有效性。
英文摘要
We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. Specifically, we consider problems in which the loss gradient is an integral with respect to a parameter-dependent distribution, a structure that includes training generative models, fine-tuning, and learning latent-variable models. We introduce mean-field dynamics and its interacting-particle approximations, which contain several existing algorithms as special cases and provides a route to constructing new methods. Under well-posedness and joint contractivity assumptions, we prove exponential convergence and show that the continuous-time particle system admits a non-asymptotic error bound. We illustrate it by developing momentum and higher-order Langevin variants and evaluating them on maximum marginal-likelihood estimation and energy-based-model training.