解耦频域感知扩散模型中的双图像参考以实现个性化生成
Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
- Hefei University of Technology(合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对个性化生成中文本与图像不对齐问题,提出Dual-FDM,通过频域掩码解耦双参考,分别处理定制与颜色风格迁移,实验证明其优于现有扩散模型。
AI中文摘要:
个性化图像生成旨在根据参考图像合成文本驱动的图像,主要将生成过程视为前景的图像定制和背景的风格迁移。先前的扩散模型方法在去噪过程中,图像定制存在与背景的文本不对齐问题,而风格迁移则存在与前景的文本不对齐问题。我们观察到,这些问题的根源在于去噪过程中混合频带之间的纠缠。为解决这一显著局限,本文研究了基于双参考(定制参考和颜色及风格参考)的个性化生成,并提出了一种范式,在频域感知扩散模型(称为Dual-FDM)中解耦这些双图像参考,通过频域内的掩码策略解耦不同频带,同时处理两个关键的个性化图像生成任务:定制风格迁移和颜色风格迁移。对于定制风格迁移,我们用定制参考前景的中频带替换风格参考中背景的中频带。对于颜色风格迁移,我们用颜色参考前景和背景的低频带替换风格参考中背景的低频带。替换后的频带用作键和值,以重建去噪个性化图像的前景和背景。大量实验验证了Dual-FDM在个性化图像生成方面优于最先进的扩散模型。我们的代码可从以下网址访问:https URL。
英文摘要:
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image.Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from https://github.com/htyjers/Dual-FDM.