AI 中文总结
本文从连续动力学出发,统一推导了HMC等基于梯度的采样器,分析了影响采样效率的几何设计选择,给出了提升采样性能的实用方法,兼具教程性与实用性。
AI 中文摘要
基于梯度的马尔可夫链蒙特卡洛(MCMC)方法常被作为算法目录介绍,包括哈密顿量蒙特卡洛(HMC)、梅特罗波利斯调整朗之万算法(MALA)、无回跳采样器(NUTS)及多个欠阻尼变体。这种呈现方式掩盖了这些方法的共同结构,更重要的是掩盖了原则上正确的采样器在实践中可能无效的原因。本文从代表理想化采样方法的精确连续时间动力学出发,这类动力学无需梅特罗波利斯调整;数值离散化使动力学在计算上可行但引入偏差,梅特罗波利斯调整通过将数值误差转化为拒绝操作消除渐近偏差,从而得到HMC、MALA、NUTS及梅特罗波利斯调整动力学朗之万算法(MAKLA)。论文后半部分介绍决定实用性能的几何设计选择:尽管MAKLA和NUTS具有良好的理论性质,其采样效率在实践中可能较低;固定质量矩阵可全局白化各向异性目标,常能解决大数据贝叶斯后验中的采样低效问题;而分层后验会引入自身问题,导致海森矩阵的状态依赖变化(如Neal漏斗),本文解释如何使用随机步长有效从这类分布中采样。本文既是基于梯度采样机制的教程,也是提高采样器性能的实用方法集。
英文摘要
Gradient-based Markov chain Monte Carlo methods are often introduced as a catalog of algorithms: Hamiltonian Monte Carlo (HMC), the Metropolis-adjusted Langevin algorithm (MALA), the No-U-Turn Sampler (NUTS), and several underdamped variants. This presentation obscures the common structure of the methods and, more importantly, the reasons why a sampler that is correct in principle may be ineffective in practice. We develop a unified account, beginning with exact continuous-time dynamics that represent idealized sampling methods and for which Metropolis adjustments are not required. Numerical discretization makes the dynamics computationally feasible but introduces bias. Metropolis adjustment removes the asymptotic bias by converting numerical errors into rejection, leading to HMC, MALA, NUTS, and the Metropolis-adjusted kinetic Langevin algorithm (MAKLA). The second half of the paper presents geometric design choices that determine practical performance, namely, although MAKLA and NUTS have nice theoretical properties, their sampling efficiency may be slow in practice. Importantly, a fixed mass matrix can whiten globally anisotropic targets, often fixing sampling inefficiency in Bayesian posteriors with large data. Whereas hierarchical posteriors introduce their own problem, causing state-dependent variation in the Hessian (e.g., Neal's funnel). We explain how a randomized step size can be used effectively to sample from such a distribution. The resulting paper is both a tutorial on the mechanics of gradient-based sampling and a set of practical recipes to improve sampler performance.