AI 中文总结
针对文本到图像扩散模型多概念生成时的遗漏或合并错误,提出无训练框架RTD,通过SOD和IGR校正初始概念分配,在AE-Bench上实现组合保真度与效率的显著提升
AI 中文摘要
文本到图像扩散模型可良好生成单个概念,但面对多个概念时,常出现概念遗漏或错误合并。我们将这些失败追溯至早期协调瓶颈:去噪开始前,提示条件注意力可能将不同概念分配到高度重叠的空间支撑区域,随着去噪推进,会使这些概念的注意力保持耦合。该观察促使我们将组合生成视为边界条件问题,而非反复控制演化轨迹。为此,我们提出无训练框架Rectify-then-Diffuse(RTD),在标准去噪前仅校正一次初始分配。首先,我们提出Soft-Overlap Disentanglement(SOD),将试点概念图之间的归一化重叠转换为可微且与布局无关的分离目标。其次,我们引入Isotropic Gradient Rectification(IGR),该方法对SOD梯度进行归一化,并应用跨提示和初始化具有一致尺度的有界潜在位移。大量实验表明,RTD实现了最先进的组合保真度和稳健提升。在AE-Bench对象对子集上,RTD相比CO3将BLIP-VQA提升45.8%,ImageReward提升19.6%,同时运行速度快2.3倍。代码将发布在该https URL
英文摘要
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3$\times$ faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse