arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17750math.NAcs.NAmath.PRstat.CO

随机梯度Langevin动力学的计算成本研究

On the computational cost of Stochastic Gradient Langevin Dynamics

Mateusz B. Majka, Tigran Nagapetyan, Łukasz Szpruch, Yue Wu, Danqi Zhuang

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究了随机梯度Langevin动力学(SGLD)与Euler-Maruyama方法在有限和漂移随机微分方程中的计算成本权衡,推导了复杂度估计,指出SGLD在大多数参数区域更优,并在小数据、激进子采样区域EM更优,转变尺度为m与ε^{-1}成正比。

中文摘要 AI 辅助

随机梯度Langevin动力学(SGLD)通过用小批量近似替代全数据集漂移评估来降低基于Langevin采样的成本,但由此产生的子采样误差可能抵消这种计算节省。我们针对具有有限和漂移的随机微分方程研究这一权衡,并将SGLD与Euler-Maruyama(EM)方法的计算成本进行比较。对于给定的均方精度$\varepsilon^2$,我们推导出复杂度估计,明确追踪其对数据集大小$m$、小批量大小$s$和精度参数$\varepsilon$的依赖关系。所得比较揭示了两种方法各自更优的不同参数区域。特别是,EM仅在数据量小、激进子采样的区域具有较低的首阶成本,而SGLD在其余大部分参数空间中更受青睐。在实际相关区域$s \ll m$中,两种方法之间的转变发生在尺度$m \asymp \varepsilon^{-1}$处。我们通过基于高斯贝叶斯推断模型的数值实验补充理论分析,检验了预测的成本区域以及潜在的离散误差和方差估计。

英文摘要

Stochastic Gradient Langevin Dynamics (SGLD) reduces the cost of Langevin-based sampling by replacing full-dataset drift evaluations with mini-batch approximations, but the resulting subsampling error may offset this computational saving. We study this trade-off for stochastic differential equations with finite-sum drifts and compare the computational cost of SGLD with that of the Euler-Maruyama (EM) method. For a prescribed mean-square accuracy $\varepsilon^2$, we derive complexity estimates that explicitly track the dependence on the dataset size $m$, mini-batch size $s$, and accuracy parameter $\varepsilon$. The resulting comparison reveals distinct parameter regimes in which either method is preferable. In particular, EM can have lower leading-order cost only in a small-data, aggressive-subsampling regime, whereas SGLD is favoured over most of the remaining parameter space. In the practically relevant regime $s \ll m$, the transition between the two methods occurs at the scale $m \asymp \varepsilon^{-1}$. We complement the theoretical analysis with numerical experiments based on a Gaussian Bayesian inference model, which examine the predicted cost regimes together with the underlying discretisation error and variance estimates.

补充信息

↑