arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16143cs.GRcs.CVcs.MMcs.SD

AnyTalk:利用视频生成模型实现任意角色的语音动画

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

Kwan Yun, Serin Yoon, Sunjin Jung, Jung Eun Yoo, Inyup Lee, Junyong Noh

首次发表
浏览论文内容

中文总结 AI 辅助

AnyTalk是一种无需动画数据、利用视频扩散模型的新方法,可生成任意角色的3D语音动画,还能蒸馏出实时版本,降低了人工与数据需求。

中文摘要 AI 辅助

我们提出了AnyTalk,一种无需任何动画数据即可为任意角色生成3D语音动画的新方法。现有音频驱动的3D语音动画方法依赖于角色特定的训练数据或繁琐的绑定/重新网格操作,而AnyTalk通过利用在海量视频数据集上训练的最新视频扩散模型规避了这些限制。我们首先通过提出的角色特定微调(Character-specific Fine-tuning, CsF)技术,将预训练的视频扩散模型适配到目标角色。通过对3D角色的渲染图像与置零音频嵌入(代表“无动作”)进行微调,我们在保留大规模视频扩散模型动作先验的同时,消除了对动画数据的需求。随后,我们通过提出的优化过程估计 blendshape 参数,将生成的说话头部视频提升为3D语音动画。AnyTalk可在不同面部网格和blendshape配置上实现唇形同步动画,大幅减少了人工工作量和数据需求。我们还通过将AnyTalk蒸馏为精简网络AnyTalk_RT来提升可用性,从而实现实时性能。通过利用说话头部视频生成,我们的方法拓宽了音频驱动语音动画技术对任意角色的可及性。代码公开于此https URL。

英文摘要

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (\textit{CsF}) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.

发表机构

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑