发表机构
Tencent; Tongji University(腾讯; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MegaAvatar基于Wan2.2-TI2V-5B模型,通过SMPL-X 3D引导和音频交叉注意力实现可控身体头部运动、表情同步与身份保持的说话头像生成。
AI 中文摘要
本报告介绍了MegaAvatar,一个基于Wan2.2-TI2V-5B模型构建的可控说话头像生成框架。与以往主要依赖音频或参考图像条件化的说话头像方法相比,我们引入了额外的SMPL-X派生的3D引导,从而实现对身体姿态和头部运动的全局控制。具体而言,我们将驱动SMPL-X序列渲染为密集网格帧,并使用轻量级3D卷积编码器对其进行编码,其输出被注入到潜在令牌中以提供整体运动控制。此外,我们扩展了Wan2.2-TI2V-5B,增加了音频和面部交叉注意力模块,分别实现细粒度的表情控制和保持输入身份。另外,我们实现了一个音频到SMPL-X模型,根据参考图像和输入音频预测SMPL-X序列,使MegaAvatar能够支持音频驱动的推理,而无需用户提供SMPL-X帧。实验表明,MegaAvatar实现了高质量的说话头像生成,具有可控的身体和头部运动、语音同步的面部表情以及一致的身份保持。MegaAvatar还支持灵活分辨率和视频长度的推理。代码、数据集和模型将在https URL中提供。
英文摘要
This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over body pose and head motion. Specifically, we render the driving SMPL-X sequence into dense mesh frames and encode them with a lightweight 3D convolutional encoder, whose outputs are injected into the latent tokens to provide overall motion control. Furthermore, we extend Wan2.2-TI2V-5B with additional audio and face cross-attention modules to enable fine-grained expression control and preserve the input identity, respectively. In addition, we implement an audio-to-SMPL-X model to predict an SMPL-X sequence conditioned on the reference image and input audio, allowing MegaAvatar to support audio-driven inference without user-provided SMPL-X frames. Experiments show that MegaAvatar achieves high-quality talking avatar generation with controllable body and head motion, speech-synchronized facial expressions, and consistent identity preservation. MegaAvatar also supports inference with flexible resolutions and video lengths. Codes, dataset, models will be avaliable in https://github.com/Jeoyal/MegaAvatar
Comments7 pages, 6 figures