arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高保真主体驱动文本到图像生成的耦合潜在噪声引导的循环内模型适配

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

Yushun Tang, Weiming Chen, Siyi Liu, Yi Zhang, Feng Wu, Zhihai He

arXiv 2608.09244首次发表:更新:

发表机构

Southern University of Science and Technology; University of Science and Technology of China; Pengcheng Lab(南方科技大学; 中国科学技术大学; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对主体驱动文本到图像生成中参考图像变化时模型难适配且难保持主体身份的问题,提出IMA方法,通过耦合潜在噪声损失引导循环内模型适配,提升了生成性能。

AI 中文摘要

文本到图像扩散模型在根据给定文本提示生成高质量图像方面已取得显著成功。主体驱动生成旨在合成定制图像,以在文本提示指定的不同视觉语境中模仿给定参考图像中主体的外观。此处的核心挑战在于,当参考图像发生变化时,扩散模型无法高效适配不同视觉语境,同时始终保持主体身份。现有方法要么使用大型特定领域数据集训练模型,要么在实际图像生成前对参考图像进行数百次迭代的模型微调。在本研究中,我们探索了一种名为循环内模型适配(In-Loop Model Adaptation, IMA)的新方法,该方法在图像生成的实际过程中,于每一步生成时适配核心扩散模型,无需在生成过程前针对参考图像进行训练。为此,我们建立了将参考图像映射到一系列潜在变量的DDIM反演链,以及仅根据文本提示生成图像的文本到图像生成链。随后,我们引入了掩码潜在一致性损失和噪声正则化损失,以表征扩散模型与这两个链在每一步生成时的潜在噪声差异。这种耦合潜在噪声损失用于引导循环内模型适配,以保留参考图像指定的主体身份,同时保持与文本提示的准确对齐,从而实现高保真文本到图像生成。我们的大量实验表明,所提出的IMA方法显著提升了主体驱动文本到图像生成的性能。

英文摘要

Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The central challenge here is that, when the reference image changes, the diffusion model cannot efficiently adapt to different visual contexts while consistently maintaining the subject identity. Existing methods either train the model with a large domain-specific dataset or fine-tune the model using the reference image for hundreds of iterations before actual image generation. In this work, we explore a new approach, called \textit{In-Loop Model Adaptation} (IMA), which adapts the core diffusion model at each generation step during the actual process of image generation, without being trained on the reference image before the generation process. To this end, we establish a DDIM inversion chain that maps the reference image to a sequence of latent, as well as a text-to-image generation chain which generates the image from the text prompt only. We then introduce a masked latent consistency loss and a noise regularization loss to characterize the latent-noise difference between the diffusion model and these two chains at each generation step. This coupled latent-noise loss is used to guide the in-loop model adaptation to preserve the subject identity specified by the reference image while maintaining accurate alignment with the text prompt, resulting in high-fidelity text-to-image generation. Our extensive experiments demonstrate that our proposed IMA method significantly improves the performance of subject-driven text-to-image generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑