发表机构
Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种无需梯度评估的生成模型采样方法,通过高频抖动余弦替代分数项,实现能量基和扩散模型的采样,适用于黑盒时变目标的潜空间场景。
AI 中文摘要
我们提出了一种用于能量基和基于分数的生成模型的采样方法,该方法无需对模型进行梯度评估。将通常包含分数 $\nabla_\mathbf{x} \log p_\theta(\bf{x})$ 的漂移项替换为模型值的高频抖动余弦 $\sqrt{\alpha\omega}\\,\cos(\omega t + k \log p_\theta(\bf{x}))$,在高频平均极限下,可产生用于能量基模型的 Langevin 马尔可夫链蒙特卡洛和基于分数的扩散模型的逆时随机微分方程。我们证明了抖动 Itô 随机微分方程的轨迹收敛到目标随机微分方程的轨迹,由相同的布朗运动驱动,在紧致时间区间上依概率一致收敛,通过将有界极值搜索(ES)扩展到 Itô 过程的平均化论证,在全局有界条件下具有显式的 $O(\omega^{-1/2})$ 均方速率。该方法不局限于光滑目标:它扩展到具有不连续曲率的 $C^{1,1}$ 能量(无需椭圆性要求)以及 Hessian 仅在零测度集外存在的 Sobolev 能量;对于具有梯度扭结的 Lipschitz 能量,平均化极限仍然适定;Krylov-Röckner 可积类是证明性的边界。该方法对每步更新速率提供了硬性先验界限,并适用于有限时间范围内的显式时变目标。在实用的评估预算下,无梯度像素空间采样无法与经过良好调优的基于反向传播的采样器竞争;该方法具有优势的场景是当模型为黑盒且目标随时间漂移时的潜空间采样。我们展示了在 CelebA-HQ($256{\times}256$)上从有限的 1D 投影测量中对时变图像进行潜空间跟踪,以及在 CIFAR-10 上进行潜空间 EBM 采样。
英文摘要
We introduce a sampling approach for energy- and score-based generative models that requires no gradient evaluations of the model. Replacing the drift term that would normally contain the score $\nabla_\mathbf{x} \log p_θ(\bf{x})$ with a high-frequency dithered cosine of the model's \textit{value}, $\sqrt{αω}\,\cos(ωt + k \log p_θ(\bf{x}))$, produces, in the high-frequency averaging limit, Langevin Markov chain Monte Carlo for energy-based models and the reverse-time SDE of score-based diffusion. We prove that trajectories of the dithered Itô SDE converge to those of the target SDE, driven by the same Brownian motion, uniformly on compact time intervals in probability, by an averaging argument that extends bounded extremum seeking (ES) to Itô processes, with an explicit $O(ω^{-1/2})$ mean-square rate under global bounds. The approach is not confined to smooth targets: it extends to $C^{1,1}$ energies with discontinuous curvature (without ellipticity requirement) and to Sobolev energies whose Hessians exist only off measure zero sets; for Lipschitz energies with gradient kinks the averaged limit remains well posed; the Krylov-Röckner integrability class is the boundary of provability. The approach provides a hard \textit{a priori} bound on the per-step update rate and applies to explicitly time-varying targets on finite horizons. Gradient-free pixel-space sampling is not competitive with well-tuned backpropagation-based samplers at practical evaluation budgets; the regime where the approach offers an advantage is latent-space sampling when the model is a black box and the target drifts in time. We demonstrate latent-space tracking for time-varying images on CelebA-HQ ($256{\times}256$) from limited 1D projection measurements and latent-space EBM sampling on CIFAR-10.