发表机构
LMU Munich; Konrad Zuse School of Excellence in Reliable AI (relAI); Munich Center for Machine Learning (MCML)(慕尼黑大学; 康拉德·祖斯可靠人工智能卓越学院; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像扩散模型缺乏特定概念引导和局部一致性可靠性不足的问题,提出无训练的CoG方法,利用概念互信息强化相关层影响,在PixArt-alpha等模型上实现性能提升。
AI 中文摘要
文本到图像扩散模型存在两个严重限制其实用性的主要缺陷:(1)标准模型缺乏内在的、针对特定概念的连续引导机制(例如无法精确控制图像的美观程度);(2)在需要高局部一致性的任务(例如生成文字或人手)上可靠性不足。为解决这些问题,我们引入了一种新颖的概念互信息概念,发现不同层之间存在与概念相关的巨大差异,证明特定结构的生成定位于网络的不同部分。我们利用这一见解,在概念引导(Concept Guidance,CoG)中强化与概念相关层的影响,CoG是一种精确的、针对目标的引导方法,可直接使用现有模型,无需额外训练、外部模型、梯度或提示工程。CoG首先量化每个层的概念特定影响,然后通过跳过与概念相关层生成的预测的加权组合来引导去噪过程。我们在各种目标及流行模型(如PixArt-alpha、SD3、SD3.5和FLUX.1-dev)上验证了性能提升,代码可在该https URL获取。
英文摘要
Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.g., for precisely controlling how aesthetically pleasing an image looks), and (2) they lack reliability for tasks requiring high local coherence (e.g., generating text or human hands). To tackle these issues, we introduce a novel notion of concept-wise mutual information and find large, concept-dependent differences between individual layers, demonstrating that the generation of specific structures is localized in distinct parts of the network. We exploit this insight by reinforcing the impact of concept-relevant layers in Concept Guidance (CoG), a precise, target-specific guidance method that works for models out-of-the-box without additional training, external models, gradients, or prompt engineering. CoG first quantifies each layer's concept-specific impact and then guides denoising using a weighted combination of predictions generated with concept-relevant layers skipped. We demonstrate performance increases across various targets and popular models like PixArt-alpha, SD3, SD3.5, and FLUX.1-dev. Code is available at https://github.com/CompVis/concept_guidance
CommentsAccepted at GCPR 2026 (Oral). 28 pages, includes supplementary material