AI 中文总结
研究针对现有基于扩散的文本到图像生成方法控制条件单一的问题,提出MixDiffusion框架,通过集成多个预训练单条件模型,支持多条件图像合成,具有无需训练、易部署和可扩展新模态的特点。
AI 中文摘要
文本到图像(T2I)生成的最新进展通过纳入文本之外的条件实现了可控图像合成。然而,大多数现有的基于扩散的方法仅限于单一类型的控制条件,限制了其灵活性。为解决此限制,我们提出MixDiffusion,一种用于多条件T2I生成的无需训练的扩散框架。MixDiffusion理论上支持任意数量的控制条件,通过协作集成多个预训练的单条件扩散模型。其关键在于通过推导的积分公式从多个单条件模型的预测噪声分布中得出多条件图像生成模型每个去噪步骤的预测噪声分布,且有严格理论证明。因其无需训练,易于部署并可扩展到新的控制模态。
英文摘要
Recent advances in text-to-image (T2I) generation have enabled controllable image synthesis by incorporating conditions beyond text. However, most existing diffusion-based methods are limited to a single type of control condition (e.g., bounding boxes or keypoints), which restricts their flexibility. To address this limitation, we propose MixDiffusion, a training-free diffusion framework for multi-condition T2I generation. MixDiffusion theoretically supports an arbitrary number of control conditions, including bounding boxes, keypoints, sketches, depth maps, reference images, and text, by collaboratively integrating multiple pre-trained uni-condition diffusion models. The key insight of the proposed approach is to derive the predicted noise distribution in each denoising step of the diffusion-based multi-condition image generation model from the predicted noise distributions of multiple diffusion-based uni-condition models with a derived integration formula, which is supported by rigorous theory proof. Owing to its training-free nature, MixDiffusion is easy to deploy and readily extensible to new control modalities.
Comments15 pages, 8 figures