arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2403.06363cs.CV

任意内容任意风格

Say Anything with Any Style

  • Shanghai Jiao Tong University(上海交通大学)
  • Netease Fuxi AI Lab(网易伏羲AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shuai Tan, Bin Ji, Yu Ding, Ye Pan

更新

AI总结:

提出SAAS动态权重方法,利用风格码本和生成模型提取说话风格,结合残差架构与HyperStyle生成风格化说话头,在唇形同步和风格化表情上超越现有方法,并扩展至视频风格编辑。

AI中文摘要:

生成具有多样头部动作的风格化说话头对于实现自然逼真的视频至关重要,但仍然具有挑战性。以往的工作要么采用回归方法捕捉说话风格,导致风格粗糙且在所有训练数据上平均化;要么使用通用网络合成不同风格的视频,导致性能次优。为解决这些问题,我们提出了一种新颖的动态权重方法,即任意内容任意风格(SAAS),该方法通过带有学习风格码本的生成模型查询离散风格表示。具体而言,我们开发了一个多任务VQ-VAE,结合三个紧密相关的任务来学习风格码本作为风格提取的先验。该离散先验与生成模型一起,增强了从给定风格片段中提取说话风格的精度和鲁棒性。通过利用提取的风格,采用由规范分支和风格特定分支组成的残差架构,根据任意驱动音频预测嘴形,同时将说话风格从源转移到任何期望的目标。为了适应不同的说话风格,我们避免使用通用网络,而是探索一种精细的HyperStyle来为风格分支生成风格特定的权重偏移。此外,我们构建了一个姿态生成器和一个姿态码本来存储量化的姿态表示,从而能够采样与音频和提取风格对齐的多样头部动作。实验表明,我们的方法在唇形同步和风格化表情方面均超越了最先进的方法。此外,我们将SAAS扩展到视频驱动的风格编辑领域,并取得了令人满意的性能。

英文摘要:

Generating stylized talking head with diverse head motions is crucial for achieving natural-looking videos but still remains challenging. Previous works either adopt a regressive method to capture the speaking style, resulting in a coarse style that is averaged across all training data, or employ a universal network to synthesize videos with different styles which causes suboptimal performance. To address these, we propose a novel dynamic-weight method, namely Say Anything withAny Style (SAAS), which queries the discrete style representation via a generative model with a learned style codebook. Specifically, we develop a multi-task VQ-VAE that incorporates three closely related tasks to learn a style codebook as a prior for style extraction. This discrete prior, along with the generative model, enhances the precision and robustness when extracting the speaking styles of the given style clips. By utilizing the extracted style, a residual architecture comprising a canonical branch and style-specific branch is employed to predict the mouth shapes conditioned on any driving audio while transferring the speaking style from the source to any desired one. To adapt to different speaking styles, we steer clear of employing a universal network by exploring an elaborate HyperStyle to produce the style-specific weights offset for the style branch. Furthermore, we construct a pose generator and a pose codebook to store the quantized pose representation, allowing us to sample diverse head motions aligned with the audio and the extracted style. Experiments demonstrate that our approach surpasses state-of-theart methods in terms of both lip-synchronization and stylized expression. Besides, we extend our SAAS to video-driven style editing field and achieve satisfactory performance.

补充信息

↑