arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PixReenact:用于流式头部化身重演的像素条件因果视频扩散

PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment

Gavriel Habib, Dvir Samuel, Or Shimshi, Rami Ben-Ari

arXiv 2610.05233首次发表:更新:

发表机构

OriginAI; NVIDIA(OriginAI; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PixReenact提出像素条件因果视频扩散框架,无需专门身份或运动表示,实现低延迟流式头部化身重演,在跨身份基准上保持身份稳定并处理极端条件。

AI 中文摘要

流式头部化身重演旨在根据实时驱动视频对参考图像进行动画化,要求稳健的运动迁移、长期的身份稳定性以及低延迟。现有方法通常依赖专门的身份或运动表示,这可能会丢弃有用的视觉信息,并继承外部提取器的失败模式。此外,许多近期基于扩散的重演方法采用离线、基于片段的方式生成,在产生输出前对整个视频片段进行联合处理和去噪,这使得连续的低延迟流式生成变得困难。我们提出PixReenact,一种基于因果视频扩散的像素条件流式重演框架。PixReenact直接以VAE编码的参考帧和驱动帧为条件,无需专门的身份或运动表示。为了将参考身份与驱动运动分离,我们使用跨身份伪监督以及锚定到原始参考和驱动输入的校正目标进行训练。长时间的自回归展开减少了自回归漂移,而状态感知的双教师蒸馏分别处理冷启动和稳态生成。在三个跨身份基准和一个长时程流式基准上,PixReenact展示了稳健的跨身份重演,特别是在极端视角、遮挡和夸张面部表情等挑战性条件下,同时在长流中保持参考身份。一个4-NFE滚动学生每次更新连续发射四帧,平均发射延迟为239毫秒。

英文摘要

Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑