arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于身份保留视频生成的表情多样化参考资料

Expression-Diverse References for Identity-Preserving Video Generation

Tianwen Fu, Wenbin Teng, Gonglin Chen, Junyi Ouyang, Haolin Xiong, Yajie Zhao

arXiv 2610.11023首次发表:更新:

发表机构

Institute for Creative Technologies, University of Southern California(南加州大学创意技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对身份保留视频生成中表情变化导致的身份表征不足问题,构建表情多样化参考集并扩展Stand-In模型,在受控基准上显著提升了不同表情强度下的身份相似度。

AI 中文摘要

身份保留视频生成旨在合成真实视频的同时保持主体的身份。然而,单张参考肖像仅捕捉到主体在一种面部构型下的外观,当表情变化时,面部外观会以高度身份特异性的方式变化,仅靠参考资料无法确定主体在未见过的表情下的外观。这种依赖表情的变化也使评估变得复杂:即使是同一个人的真实图像,在强烈表情下与中性参考资料的相似度也可能降低。我们从生成和评估两个角度研究这一局限。首先,我们使用受控照片和MEAD视频量化人脸识别相似度如何随表情强度变化。然后,我们构建一个紧凑但具有表现力的参考图库,以捕捉不同的依赖表情的面部构型,与该图库匹配能为表情动作下的身份相似度提供更稳健的度量。为进一步揭示随表情强度增加的性能下降,我们分别报告轻度、强烈和极端表情下的身份相似度。在生成方面,我们扩展Stand-In以基于我们的表情多样化参考集进行条件生成,并开发了一个数据整理流程,从训练视频中提取一致且多样化的面部裁剪图。在仅可获得单张肖像的实际场景中,我们使用预训练的面部重演模型合成额外表情来构建参考集。在我们的受控基准上,真实参考集和合成参考集在所有三种表情强度 regime 下的身份相似度均优于所评估的基线,其中极端表情下的改进最大。

英文摘要

Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait captures the subject's appearance under only one facial configuration. As expressions change, facial appearance can vary in highly identity-specific ways, leaving the subject's appearance under unseen expressions underdetermined by the reference alone. This expression-dependent variation also complicates evaluation: similarity to a neutral reference may decrease under strong expressions even for real images of the same person. We investigate this limitation from both generation and evaluation perspectives. First, we quantify how face-recognition similarity varies with expression intensity using controlled photographs and MEAD videos. We then construct a compact yet expressive reference gallery that captures diverse expression-dependent facial configurations. Matching against this gallery provides a more robust measure of identity similarity under expressive motion. To further expose performance degradation with expression intensity, we report identity similarity separately for mild, intense, and extreme expressions. For generation, we extend Stand-In to condition on our expression-diverse reference sets and develop a data-curation pipeline that extracts consistent yet diverse face crops from training videos. In practical settings where only a single portrait is available, we construct the reference set by synthesizing additional expressions with a pretrained facial reenactment model. On our controlled benchmark, both real and synthesized reference sets outperform the evaluated baselines in identity similarity across all three expression-intensity regimes, with the largest improvements for extreme expressions.

Comments9 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑