发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对合成图像微调导致主体保真度下降的问题,提出无需训练的采样时校正方法 ReGain,通过按频段缩放引导来消除 CFG 膨胀,显著缩小保真度差距并保持文本对齐。
AI 中文摘要
文本到图像扩散模型通过 DreamBooth 在少量主体图像上进行微调,从而实现对特定主体的个性化。如今,这些图像越来越多地来自扩散模型而非相机。我们表明,在此类合成图像上进行微调会降低主体保真度,产生过饱和的颜色和过多的高频细节。为隔离原因,我们从同一基础模型出发,使用相同的 DreamBooth 配方微调两个模型:一个在主体的真实照片上微调,另一个在由前者生成的该主体的合成图像上微调。我们将退化追溯到无分类器引导(CFG)。对于在合成图像上个性化的模型,条件噪声预测与无条件噪声预测之间的夹角,以及它们差值的范数,远大于在真实照片上个性化的模型。这种膨胀向高频方向增长,并且也出现在与主体语义接近的其他提示词(如其类别名词)上,但不会出现在无关提示词上。我们提出 ReGain,一种无需训练的采样时校正方法,它测量每个频段的引导相对于基础模型的膨胀程度,并相应缩小该频段。ReGain 不需要真实照片。在 Stable Diffusion v1.5 上,ReGain 缩小了与真实照片个性化模型之间主体保真度差距的 51-64%,这一结果由 DINO、DINOv2 和 CLIP-I 衡量。它还在 SDXL 和 SD 3.5 上提高了主体保真度,并在所有三个骨干网络上保持了文本对齐。
英文摘要
Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.