看见语音:学习可见发音动态用于语音驱动的3D面部动画
Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出发音感知框架,通过方向性发音运动与记忆检索组合,实现语音驱动的3D面部动画,在标准指标和用户研究中均达最优。
AI中文摘要:
近期语音驱动的3D面部动画研究进展提升了顶点级重建质量,但语音一致的可见发音仍具挑战性。这是因为语音产生遵循结构化和受约束的发音器官协调,且从声学到运动的映射本质上是一对多的。受可见发音结构化模式的启发,我们提出了一种新颖的发音感知框架,通过方向性发音运动建模可见语音,并将其组合成表面一致的3D面部运动。为用三种方向性发音运动(展开、张开和突出)表示可见发音,我们提出了语音-发音记忆(SAM),通过基于键值记忆结构的检索和解码,在语音语境下捕获语音与这些运动之间的对应关系。随后,拓扑感知发音组合(TAC)在网格拓扑下整合预测的方向性发音运动,以生成表面一致的3D面部运动。在VOCASET和TFHP上的实验表明,我们的方法在标准重建指标上达到最先进性能,并改善了唇部发音的可见发音距离和速度误差,而用户研究确认了在唇形同步和真实感方面的明显偏好。
英文摘要:
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.