arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散模型的学习端到端引导调度

Learned End-to-End Guidance Schedules for Diffusion Models

Aneesh Barthakur, Mathias Niepert, Luiz F. O. Chamon

arXiv 2610.01502首次发表:更新:

发表机构

University of Stuttgart; École polytechnique, Institut Polytechnique de Paris(斯图加特大学; 巴黎综合理工学院,巴黎综合理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LEEGS方法,通过学习时间依赖的引导调度,以更少的采样步骤实现扩散模型在数据质量与需求满足间的平衡,并在多种引导任务上以更低计算成本达到或超越基线性能。

AI 中文摘要

扩散模型是一种强大的生成范式,广泛应用于多媒体和科学应用。引导扩散方法通过在推理过程中将可微损失(引导函数)的梯度作为漂移项添加,从而对生成过程施加要求。该漂移的权重(引导尺度)对于数据质量与需求满足之间的权衡至关重要。为了同时实现这两个目标,引导扩散必须采用较小的引导尺度和较长的采样过程,这导致高昂的计算成本。本文提出学习端到端引导调度(LEEGS),以更少的采样步骤实现这些目标。LEEGS通过使用随机梯度下降在一小组示例上最小化引导函数来训练一个时间依赖的调度。通过引导采样进行反向传播在计算上代价高昂,因此LEEGS使用梯度的近似值,将训练时间缩短了4倍。我们在多种引导任务上评估LEEGS,包括(a)图像修复,(b)噪声图像逆问题,(c)人脸身份引导生成,以及(d)正向和逆向偏微分方程问题,在相同预算(50或100次神经函数评估)下优于基线,或仅用10%的步骤匹配恒定引导的效果。

英文摘要

Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, (c) face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑