arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24124cs.CV

ViDS:使用3D面部跟踪的视频扩散着色器

ViDS: Video Diffusion Shader using 3D Face Tracking

Wenbo Ji, Davide Davoli, Zhe Chen, Liam Schoneveld, Matthias Nießner, Jiapeng Tang

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用3D面部跟踪实现肖像动画,先重建3DMM网格,再用驱动视频参数动画处理,借助视频扩散模型合成动画,引入自回归扩散采样过程,相比之前方法能更精细控制表情姿势,且保留身份外观,消融研究验证了有效性。

中文摘要 AI 辅助

我们介绍了ViDS,一种利用3D面部跟踪实现富有表现力且保持身份的肖像动画的视频扩散着色器。首先从参考图像重建特定身份的3DMM网格,然后使用驱动视频中的表情和姿势参数对其进行动画处理。利用3DMM法线贴图的密集几何线索,采用视频扩散模型作为神经着色器,在保留参考图像外观和身份的同时合成逼真的肖像动画。更准确的3DMM跟踪能实现更精细的表情控制。还引入自回归扩散采样过程,扩展生成范围并减少相邻片段间的不连续性。与先前基于扩散的肖像动画方法相比,该方法在忠实保留身份和外观的同时,实现了更详细、一致的表情和姿势控制。详细的消融研究验证了设计选择的有效性。

英文摘要

We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstruct the identity-specific 3DMM mesh from the reference image, and then animate it using expression and pose parameters from a driving video. Leveraging dense geometric cues from 3DMM normal maps, we employ a video diffusion model as a neural shader to synthesize lifelike portrait animations while preserving the appearance and identity of the reference image. We find that more accurate 3DMM tracking enables finer-grained expression control. We also introduce an autoregressive diffusion sampling process that extends generation beyond the model's native window while reducing discontinuities between adjacent clips. Compared with prior diffusion-based approaches for portrait animation that rely on landmark-based conditioning or implicit motion latents, our method achieves more detailed and consistent expression and pose control while faithfully preserving identity and appearance. Detailed ablation studies validate the effectiveness of our design choices. Project page: https://fusheng-ji.github.io/ViDS/

发表机构

  • Technical University of Munich(慕尼黑工业大学)
  • Toyota Motor Europe NV/SA(丰田汽车欧洲股份公司)
  • Woven by Toyota(丰田编织)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑