发表机构
LIVIA, École de technologie supérieure; CentraleSupélec, Université Paris-Saclay; Inria, Université Côte d’Azur(LIVIA,蒙特利尔高等技术学院; 巴黎中央理工-高等电力学院,巴黎-萨克雷大学; 法国国家信息与自动化研究所,蔚蓝海岸大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BEACON通过解耦视觉身份与表达行为,利用参考图像和参考视频双信号条件生成,在MEAD和RAVDESS上以少量微调提升面部表现力并保持身份。
AI 中文摘要
生成既能保持视觉身份又能保留个体特有表达行为的人类中心视频仍是一项基本挑战。除了再现外观,模型还必须复制那些表征主体随时间表达情感的面部行为。然而,大多数最先进的方法仅基于单张参考图像进行条件生成,而该图像不包含任何关于这些时间动态的信息。因此,它们往往能保留主体的视觉身份,但产生的表情变化有限且主体特异性较弱。为缓解这一问题,我们提出了BEACON,一个用于主体特定视频生成的轻量级框架,通过将视觉身份与表达行为解耦来生成更具表现力的视频。BEACON基于两种互补信号进行条件生成:一个编码身份的参考图像和一个捕捉主体特定面部动态的参考视频。通过基于这些互补信号进行条件化,BEACON生成的视频能更好地同时保留主体的外观和特征性面部动态,同时支持身份-表情迁移。我们在MEAD和RAVDESS数据集上的实验表明,通过在大约2000对数据上微调并更新预训练的Wan视频扩散模型约1%的参数,BEACON在面部表现力上超越了最先进的视频生成方法,同时保持了具有竞争力的身份保持能力。
英文摘要
Generating human-centric videos that preserve both visual identity and person-specific expressive behavior remains a fundamental challenge. In addition to reproducing appearance, a model must replicate the facial behaviors that characterize how a subject expresses emotion over time. However, most state-of-the-art methods condition generation on a single reference image, which contains no information about these temporal dynamics. As a result, they tend to preserve the subject's visual identity but often produce expressions with limited variation and weak subject specificity. To mitigate this issue, we introduce BEACON, a lightweight framework for person-specific video generation that produces more expressive videos by disentangling visual identity from expressive behavior. BEACON conditions generation on two complementary signals: a reference image encoding the identity and a reference video capturing subject-specific facial dynamics. By conditioning on these complementary signals, BEACON generates videos that better preserve both the subject's appearance and characteristic facial dynamics, while also supporting identity-expression transfer. Our experiments on the MEAD and RAVDESS datasets show that by fine-tuning on approximately 2,000 pairs and updating about 1% of the pretrained Wan video diffusion model, BEACON improves facial expressivity over state-of-the-art video generation methods while maintaining competitive identity preservation.