无分类器引导调度的对抗学习
Adversarial Learning of Classifier-Free Guidance Schedules
浏览论文内容
中文总结 AI 辅助
本文提出一种对抗学习方法,将无分类器引导调度建模为密度比估计问题,训练判别器与轻量生成器,在文本到图像生成基准上优于启发式CFG调度及现有动态引导学习方法。
中文摘要 AI 辅助
现代文本到图像的扩散模型依赖无分类器引导(Classifier-Free Guidance,CFG)来实现高图像保真度和文本对齐。然而,CFG通常在所有时间步、样本和条件下应用静态全局尺度,这种选择通常次优且会引入伪影,因为不同状态可能受益于不同水平的引导。虽然时变调度已知可提升质量,但手动设计它们并非易事且依赖于应用。本文中,我们将引导调度学习为扩散时间、条件和当前噪声样本的函数,以更好地使采样图像与文本提示对齐。我们将此问题表述为密度比估计问题:训练判别器以估计真实和引导边际分布之间的时间相关对数密度比,同时轻量生成器网络预测最优的、状态相关的引导尺度。在文本到图像生成基准上的实验表明,我们的方法在启发式CFG调度和现有动态引导学习方法上表现更优。
英文摘要
Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. However, CFG typically applies a static, global scale across all timesteps, samples, and conditions -- a choice that is generally suboptimal and can introduce artifacts, as different states may benefit from different levels of guidance. While time-varying schedules are known to improve quality, designing them by hand is non-trivial and application-dependent. In this paper, we learn the guidance schedule as a function of diffusion time, conditioning and the current noisy sample, in order to better align sampled images with the text prompt. We frame this as a density ratio estimation problem: a discriminator is trained to estimate the time-dependent log-density ratio between the true and guided marginal distributions, while a lightweight generator network predicts the optimal, state-dependent guidance scale. Empirically, our approach outperforms both heuristic CFG schedules and prior methods for learning dynamic guidance on text-to-image generation benchmarks.
发表机构
- Google(谷歌公司)
- Google DeepMind(谷歌DeepMind)
- Gatsby UCL(盖茨比伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。