ETHead:从语音生成富有表现力的3D面部动画与头部运动
ETHead: Generating Expressive 3D Facial Animation and Head Movement from Speech
浏览论文内容
中文总结 AI 辅助
本文提出ETHead方法,通过自蒸馏框架与情感调制概率掩码机制,从语音生成与情感匹配的3D面部及头部运动,性能优于现有最优方法,其编码器可迁移增强其他动画框架。
中文摘要 AI 辅助
仅从语音生成富有表现力的3D说话头像仍是一项重大挑战,原因在于高保真3D数据的稀缺性,这限制了对复杂情感运动模式的建模。本文提出了ETHead(Expressive Talking Head),一种能生成与输入语音情感内容生动匹配的3D面部及头部运动的方法。为克服数据限制,我们设计了一种自蒸馏框架,利用大规模2D说话视频预训练专用语音编码器。通过引入新颖的情感调制概率掩码机制,该框架使语音表示与富有表现力的视觉动态对齐,让编码器能直接从音频中提取与面部及头部运动高度相关的特征。这些特征随后被用于指导3D生成,丰富输入线索并通过联合语音-运动潜在空间提供显式监督。大量实验表明,ETHead的性能显著优于现有最优方法。此外,我们的运动对齐语音编码器可作为可迁移模块,为增强其他3D说话头像动画框架的表现力提供通用解决方案。项目页面可访问此https URL。
英文摘要
Generating expressive 3D talking heads solely from speech remains a significant challenge due to the scarcity of high-fidelity 3D data, which limits the modeling of complex emotional motion patterns. In this paper, we introduce \textbf{E}xpressive \textbf{T}alking \textbf{Head} (ETHead), a method for generating 3D facial and head motions that vividly align with the emotional content of input speech. To overcome the data limitations, we design a self-distillation framework that leverages large-scale 2D talking videos to pre-train a specialized speech encoder. By incorporating a novel emotion-modulated probabilistic masking mechanism, this framework aligns speech representations with expressive visual dynamics, allowing the encoder to extract features highly correlated with facial and head motions directly from audio. These features are then leveraged to guide 3D generation, enriching input cues and providing explicit supervision through a joint speech-motion latent space. Extensive experiments demonstrate that ETHead substantially outperforms state-of-the-art methods. Furthermore, our motion-aligned speech encoder can serve as a transferable module, offering a general solution for enhancing expressiveness in other 3D talking head animation frameworks. The project page is available at https://verdure-oss.github.io/ETHead.github.io/.