arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05845cs.LGcs.AI

CACFG:曲率感知的无分类器引导与最优控制

CACFG: Curvature-Aware Classifier-Free Guidance and Optimal Control

Max Collins, Dasith de Silva Edirimuni, Jordan Vice, Tim French, Ajmal Mian

首次发表
浏览论文内容

中文总结 AI 辅助

针对CFG高引导强度失效问题,提出曲率感知CACFG,通过最优控制框架约束控制集,提升中高引导强度下的生成质量与多样性平衡。

中文摘要 AI 辅助

扩散模型通过学习逆转一个固定的损坏过程来生成样本,而无分类器引导(CFG)是将这一过程条件化于所需类别或提示的标准机制。CFG可以以不同的引导强度应用,虽然较高的强度能提升图像质量和条件对齐度,但过高的引导强度会降低图像质量和多样性。此外,CFG违背了原则性的扩散采样动力学,现有关于其为何在违背情况下仍有效的解释在基础理论上存在分歧,或无法扩展到实践中使用的确定性采样器。我们同时解决这两个问题。首先,我们将CFG采样构建为一个连续时间最优控制问题,将采样轨迹视为一系列控制,旨在最大化所需条件的概率。求解所得的Hamilton-Jacobi-Bellman方程表明,在使用无约束控制集时,CFG在特定路径成本下得以恢复。我们认为,缺乏约束是CFG在高引导强度下失效的原因,因为它允许采样路径任意远离当前图像估计。为解决此问题,我们提出曲率感知的CFG(CACFG),其将控制集约束在一个由训练变分自编码器时使用的高斯正则化所指导的超球面上。我们表明,CFG采样产生的控制输入经常违反此约束,并且跨扩散模型、数据集和引导调度,CACFG在中高引导强度下实现了更优的生成质量,且质量-多样性权衡不如常规CFG严重。

英文摘要

Diffusion models generate samples by learning to reverse a fixed corruption process, and classifier-free guidance (CFG) is the standard mechanism for conditioning this process on a desired class or prompt. CFG can be applied at varying guidance strengths, and while higher strengths improve image quality and conditional alignment, too high a guidance strength can degrade image quality and diversity. Furthermore, CFG violates principled diffusion sampling dynamics, and existing explanations for why it works despite the violation disagree on the underlying theory or do not extend to deterministic samplers used in practice. We address both these issues. We first frame CFG sampling as a continuous-time optimal control problem, treating the sampling trajectory as a sequence of controls chosen to maximise the probability of the desired condition. Solving the resulting Hamilton--Jacobi--Bellman equation shows that CFG is recovered under specific path costs when using an unconstrained control set. We argue this lack of constraint is responsible for CFG's failure at high guidance strengths, since it permits the sampling path to move arbitrarily far from the current image estimate. To fix this, we propose curvature-aware CFG (CACFG), which constrains the control set to a hypersphere informed by the Gaussian regularisation used when training variational autoencoders. We show that the control inputs produced by CFG sampling routinely violate this bound, and that across diffusion models, datasets, and guidance schedules, CACFG achieves superior generative quality at mid-to-high guidance strengths with a less severe quality-diversity tradeoff than regular CFG.

发表机构

  • The University of Western Australia(西澳大学)

机构由 AI 辅助整理,请以论文原文为准。

↑