用于人脸合成的身份条件潜在一致性蒸馏
Identity-Conditioned Latent Consistency Distillation for Face Synthesis
浏览论文内容
中文总结 AI 辅助
本研究针对扩散模型人脸合成计算成本高的问题,通过将Arc2Face知识蒸馏为潜在一致性模型,实现4.36倍加速且保持竞争力,相关方案已公开。
中文摘要 AI 辅助
扩散模型在高保真图像合成领域已取得优异成果,但其迭代采样过程使得大规模生成的计算成本较高。当为人脸识别任务生成合成人脸数据集时,该限制尤为突出——这类任务需要大量主体,且每个主体需具备不同姿态、表情、年龄等的多类样本。本研究表明,采用身份条件的人脸合成可通过仅需少量迭代的潜在一致性模型以大幅降低计算成本,且不会损害图像质量。训练阶段,我们将基础扩散模型Arc2Face(教师模型)的知识进行蒸馏,方法是将其原有的文本到图像流程适配为嵌入到人脸的设置,用ArcFace身份嵌入替代文本提示。我们的蒸馏模型(学生模型)生成身份条件人脸图像的平均推理时间为每张0.4819秒,而Arc2Face为每张2.102秒,速度提升达4.36倍。基于FID分数的定量结果显示,在所有评估协议中,蒸馏模型仍与Arc2Face具有竞争力。在10万张生成图像上,其在CelebA数据集上实现了近乎持平的表现(13.921 vs 12.928),在WebFace42M数据集上的表现优于教师模型(9.317 vs 9.802);对Synth-500和AgeDB的进一步评估显示,前者存在适度性能差距,后者则结果相当。这些结果表明,通过针对特定任务的潜在一致性蒸馏,可在保留大规模合成人脸生成所需的高图像质量的同时,加速Arc2Face。本研究的方案已公开,链接为this https URL。
英文摘要
Diffusion models have achieved strong results in high-fidelity image synthesis, but their iterative sampling process makes large-scale generation computationally expensive. This limitation is especially relevant when generating synthetic face datasets for face recognition, where a large number of subjects with many samples in different poses, expressions, ages, etc., are required. In this work, we show that identity-conditioned face synthesis can be performed at a substantially lower computational cost by a latent Consistency Model with few iterations, without compromising image quality. For training, we distill knowledge from the foundation Diffusion Model Arc2Face (teacher) by adapting its original text-to-image pipeline to an embedding-to-face setting, replacing textual prompts with ArcFace identity embeddings. Our distilled model (student) generates identity-conditioned face images with an average inference time of 0.4819 seconds per image, compared with 2.102 seconds for Arc2Face, resulting in a 4.36$\times$ speed-up. Quantitative results, based on FID scores, show that the distilled model remains competitive with Arc2Face across all evaluation protocols. On 100k generated images, it achieves near-parity on CelebA (13.921 vs. 12.928) and outperforms the teacher on WebFace42M (9.317 vs. 9.802). Further evaluations on Synth-500 and AgeDB show a moderate performance gap for the former but comparable results for the latter. These results indicate that Arc2Face can be accelerated through task-specific latent consistency distillation while preserving high image quality for large-scale synthetic face generation. Our proposal is publicly available at https://github.com/UFPR-IPASP-PR/FaceRec-IdentityConsistency.
发表机构
- Federal University of Paraná(巴拉那联邦大学)
- Federal Institute of Mato Grosso (IFMT)(马托格罗索联邦学院(IFMT))
机构由 AI 辅助整理,请以论文原文为准。