arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视潜在扩散模型中无分类器引导方法

Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models

Artem Sergievskii, Artyom Turevich, Sergey Kastryulin

arXiv 2608.16786首次发表:更新:

发表机构

HSE University; Yandex(高等经济大学; Yandex公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文重新评估了八种源于无分类器引导(CFG)的无训练技术,发现无一种方法始终优于CFG,APG仅获名义最佳分数,注意力扰动方法在不同模型上表现有差异,CFG仍是低成本竞争基线。

AI 中文摘要

推理时质量增强方法是提升扩散模型的有效且广泛采用的手段,无需昂贵的重新训练。本文研究了一类在概念上源于无分类器引导(CFG)的无训练技术,其中大多数最初在较旧的U-Net扩散模型上提出,且仅使用评估图像质量的指标进行验证,未考虑生成图像与其关联文本提示之间的组合对齐或语义对应。我们在两个开放权重的整流流变换器上,采用固定的每模型协议和三个组合对齐基准,重新评估了八种此类方法。没有一种方法在所有测量标准上始终优于CFG。APG获得了几个名义最佳分数,但相应的增益通常仍在估计的评估不确定性范围内。注意力扰动方法在SD3.5 Medium上提供了孤立的增益,而在FLUX.2 [klein] 4B Base上则更频繁地导致性能下降,同时CFG仍然是具有竞争力的低成本基线。

英文摘要

Inference-time quality-enhancement methods are an effective and widely adopted means of improving diffusion models without expensive retraining. We study a family of training-free techniques conceptually rooted in Classifier-Free Guidance (CFG), most of which were originally proposed on older U-Net diffusion models and validated using metrics that assess image quality in isolation, without accounting for compositional alignment or semantic correspondence between the generated image and its associated text prompt. We re-evaluate eight such methods on two open-weight rectified-flow transformers under a fixed per-model protocol and three compositional-alignment benchmarks. No method consistently improves on CFG across the measured criteria. APG obtains several nominal best scores, but the corresponding gains often remain within the estimated evaluation uncertainty. Attention-perturbation methods provide isolated gains on SD3.5 Medium and more frequent degradations on FLUX.2 [klein] 4B Base, while CFG remains a competitive lower-cost baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑