发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无分类器引导缺乏明确标准的问题,提出基于频谱对齐的无需训练的频谱校正引导方法,通过校正采样轨迹偏差,在文本到图像和ImageNet生成中优于基线方法。
AI 中文摘要
条件图像生成的实际成功取决于条件对齐和视觉保真度上的细微差异。无分类器引导(CFG)是这一成功的关键,但它缺乏明确的标准,使得难以评估引导轨迹是否按预期进行。为解决这一差距,我们证明频谱对齐为理解引导行为和通过自适应校正改进引导扩散采样提供了原则性标准。我们的分析将中间状态的频谱识别为与前向过程预期频谱演化一致性的指标。基于这一观察,我们引入了频谱校正引导(Spectral Correction Guidance),一种在采样过程中校正与解析参考频谱偏差的方法。所提方法无需训练,适用于各种扩散主干和条件生成任务,无需修改底层模型。实验表明,在文本到图像生成中,该方法在基于偏好的指标上持续优于基线引导方法,并在ImageNet上相比CFG提高了生成质量。这些改进在多种引导尺度下和更少的去噪步骤中依然存在。我们的分析和消融研究为引导行为及所提方法如何影响生成质量提供了见解。
英文摘要
The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progressing as intended. To address this gap, we show that spectral alignment provides a principled criterion for understanding guidance behavior and improving guided diffusion sampling through adaptive correction. Our analysis identifies the spectra of intermediate states as an indicator of consistency with the expected spectral evolution of the forward process. Based on this observation, we introduce Spectral Correction Guidance, a method that corrects deviations from an analytic reference spectrum during sampling. The proposed method is training-free and applicable across diffusion backbones and conditional generation tasks without modifying the underlying model. Experiments demonstrate consistent gains in preference-based metrics over baseline guidance methods in text-to-image generation and improved generation quality over CFG on ImageNet. These improvements persist across a range of guidance scales and with fewer denoising steps. Our analyses and ablations provide insight into guidance behavior and how the proposed method affects generation quality.