arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应二阶求解器用于快速随机扩散采样

Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling

Ella Kemperman, Luca Ambrogioni

arXiv 2610.03034首次发表:更新:

发表机构

Radboud University(拉德堡德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出自适应二阶求解器,通过比例-积分控制加速扩散采样,在图像和语言任务上以更低计算成本获得更优或相当的质量,并发现逐样本自适应的收益因问题而异。

AI 中文摘要

扩散模型依赖于需要时间离散化的数值求解器,这对采样成本与质量之间的权衡有很大影响。然而,反向过程的计算难度沿采样轨迹和跨数据分布而变化,使得离散化的选择变得重要。我们将比例-积分(PI)步长控制适应于扩散,使用我们的扩散噪声归一化误差估计器。与现有扩散中仅响应当前误差的自适应方法不同,PI求解器还结合了先前的误差,从而产生更平滑的步长适应。我们进一步表明,这些逐样本轨迹表现出共享结构,可以聚合成一个固定调度,该调度保留了自适应采样的许多好处。我们在自然图像和语言数据集上评估了这两种方法,在匹配的神经网络评估次数(NFE)下,以FID衡量的质量进行比较,并与广泛使用的随机求解器和调度进行比较。对于图像,我们的固定离散化在随机Heun采样器下,以及在低NFE下使用EDM-churn采样器时,在样本质量上优于常用的EDM调度。此外,我们的PI自适应求解器在大多数随机和自适应基线上获得了更好的FID,尽管在低NFE下它没有击败EDM-churn采样器。此外,我们发现我们的求解器在低到中等NFE下,在语言扩散上以困惑度衡量,优于EDM和熵调度,但缺点是token熵较低。最后,我们发现逐样本自适应性的好处是问题依赖的。它在1D玩具示例中非常有益,而对于图像和语言数据仅边际收益,其中平均调度有时甚至优于PI自适应求解器。代码可在https URL获取。

英文摘要

Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error estimator. Unlike existing adaptive methods in diffusion that respond only to the current error, the PI solver also incorporates the previous error, yielding smoother step adaptation. We further show that these per-sample trajectories exhibit shared structure and can be aggregated into a fixed schedule that retains much of the benefit of adaptive sampling. We evaluate both approaches on natural-image and language datasets, in terms of quality, measured by FID at a matched number of neural network evaluations (NFE), comparing them with widely used stochastic solvers and schedules. For images, our fixed discretization outperforms the commonly used EDM schedule in terms of sample quality when used with the stochastic Heun sampler, and with the EDM-churn sampler at low NFE. Additionally, our PI adaptive solver obtains better FID than most stochastic and adaptive baselines, although it does not beat the EDM-churn sampler at low NFE. Moreover, we find our solver outperforms both the EDM and the entropy schedule on language diffusion at low-to-medium NFE in terms of perplexity, with the drawback of lower token entropy. Lastly, we find that the benefit of per-sample adaptivity is problem-dependent. It is highly beneficial in 1D toy examples, while only marginal for image and language data, where the average schedule sometimes even outperforms the PI-adaptive solver. Code is available at https://github.com/ellakemperman/adaptive-second-order-diffusion-solvers

CommentsSubmitted to ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑