AI 中文总结
针对大遮挡和冲突文本引导下人脸修复的身份保留难题,提出级联扩散框架ReSem-Face,在CelebAHQ-IDI-5等数据集上验证其身份保留修复及文本编辑质量优于基线。
AI 中文摘要
基于扩散模型的人脸修复近期已取得令人印象深刻的视觉质量,但在严重遮挡和冲突文本引导下保留身份保真度仍是重大挑战。为解决该问题,我们提出人脸参考语义修复模型(ReSem-Face),这是一种级联扩散框架,引入显式身份条件语义先验用于多参考人脸修复。该方法从多个参考中提取代表性身份特征以重建缺失语义区域,随后通过多流条件架构引导扩散过程。此设计在像素缺失时提供强语义约束,稳定身份重建同时兼容提示驱动编辑。在CelebAHQ-IDI-5和VGGFace2上的实验表明,ReSem-Face在严重语义掩码下能生成更可靠的身份保留修复结果,且相比代表性基线提升了文本控制编辑质量。
英文摘要
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.