arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ID-V2V:保留身份的视频风格化

ID-V2V: Identity-Preserving Video Restylization

Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu

arXiv 2607.22830首次发表:更新:

发表机构

Netflix; Eyeline Labs; Adobe(网飞公司; 眼线实验室; 奥多比公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何在视频风格化中保留人类身份和表演,提出将身份保留与视频合成解耦的方法,引入ID-V2V框架,通过互补控制信号实现,实验证明其在面部相似度等方面表现优异,优于现有方法。

AI 中文摘要

在视觉叙事中,人类表演对创作意图和叙事意义至关重要。然而,对于生成式视频模型来说,在进行灵活视觉编辑时保留人类身份和表演仍然具有挑战性。我们将此挑战形式化为保留身份的视频风格化,即在保留面部相似度和表演的同时,将编辑关键帧指定的场景、灯光和风格变化传播到源视频。一个关键障碍是缺乏配对训练数据。为解决此问题,我们提出将基于源的身份保留与编辑驱动的视频合成解耦。我们将身份保留视为视频重新打光问题,将视觉编辑传播建模为由编辑关键帧引导的受控视频合成。在此基础上,我们引入了ID-V2V,一个集成互补控制信号的视频到视频生成框架。大量实验表明,ID-V2V在保留面部相似度和细粒度面部表演方面显著优于现有方法,支持单主体和多主体场景,并具有高视觉质量。

英文摘要

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.

CommentsAccepted to SIGGRAPH Asia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑