Diff-ID:通过扩散模型实现身份一致的面部图像生成与变形
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
- University of Surrey(萨里大学)
- Jiangnan University(江南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对面部图像合成中高分辨率下身份保持难题,提出Diff-ID框架,通过自定义数据集、集成嵌入及新损失函数实现,实验表明其在身份-真实感权衡上表现出色,还展示了基于DDIM的变形管道,强调联合评估身份保持与真实感。
AI中文摘要:
生成扩散模型革新了面部图像合成,但在高分辨率输出中稳健保持身份仍是关键挑战,对安全系统等应用至关重要。我们引入Diff-ID,这是一个基于扩散的框架,通过自定义210K图像数据集及微调BLIP模型增强身份感知,在微调的Stable Diffusion UNet中集成ArcFace和CLIP嵌入,还提出基于ArcFace余弦相似度的伪判别器损失。实验表明Diff-ID在面部相似度上不超InstantID,但FID更低,FIQ最佳。我们还展示了基于DDIM的变形管道,强调应联合评估身份保持和真实感,报告了结合身份相似度和感知真实感的Face Image Quality (FIQ)分数。
英文摘要:
Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is especially vital for security systems, biometric authentication, and privacy sensitive applications, where any drift in identity integrity can undermine trust and functionality. We introduce Diff-ID, a diffusion based framework that enforces identity consistency while delivering photorealistic quality. Central to our approach is a custom 210K image dataset synthesized from CelebA-HQ, FFHQ, and LAION-Face and captioned via a fine tuned BLIP model to bolster identity awareness during training. Diff-ID integrates ArcFace and CLIP embeddings through a dual cross attention adapter within a fine tuned Stable Diffusion UNet. To further reinforce identity fidelity, we propose a pseudo discriminator loss based on ArcFace cosine similarity with exponential timestep weighting. Experiments on held out and unseen faces show that Diff-ID does not exceed InstantID in raw ArcFace Face Similarity, but achieves substantially lower FID and the strongest FIQ based identity--realism trade off among the evaluated methods. We also present a unified DDIM based morphing pipeline that enables qualitative facial interpolation without per identity fine tuning. We further argue that identity preservation and photorealism should be evaluated jointly rather than in isolation, as high identity similarity alone does not guarantee realistic outputs. To make this trade off explicit, we report Face Image Quality (FIQ) as a complementary ratio based score that combines identity similarity and perceptual realism while keeping FS and FID as the primary metrics.