发表机构
Shanghai Jiao Tong University; Tencent Games(上海交通大学; 腾讯游戏)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SubtleTalk通过多条件建模和残差流匹配,结合新构建的大规模数据集SubtleTalk-Face,提升3D Talking Heads弱相关面部动力学的真实感与多样性,同时保留准确唇同步。
AI 中文摘要
音频驱动的3D面部动画旨在从语音中合成逼真且时间连贯的动作。尽管唇同步技术已取得显著进展,但对于照片级真实感面部动画至关重要的弱相关动力学(包括眉毛运动、眨眼和头部运动)仍难以被忠实地建模,且常呈现静态或不自然的重复状态。我们将这一局限性归因于三个因素:(a) 弱相关动力学的条件信息不足;(b) 确定性回归捕捉多样化运动模式的能力有限;(c) 来自不可靠的上半身伪标签及有限数据集多样性的数据瓶颈。为解决这些问题,我们提出SubtleTalk,这是一个通过多条件建模和残差流匹配生成自然且可控的弱相关面部动力学的框架。首先,为弥补仅语音指导的不足,我们引入可解释控制项,包括韵律、区域强度和效价-唤醒(Valence-Arousal)信号,以明确捕捉弱相关动力学的时间、幅度和情感变化。其次,为克服确定性回归表达能力有限的问题,我们基于稳定的语音驱动运动先验构建残差流匹配,使模型能够捕捉确定性预测之外的随机偏差。第三,为缓解数据瓶颈,我们构建SubtleTalk-Face,这是一个包含约3900个身份和74小时数据的大规模3D面部动画数据集,通过简单且可扩展的伪标注流程构建,具备改进的上半身跟踪和帧级VA标注。大量实验表明,我们的方法在保持准确唇同步的同时,显著提升了弱相关面部动力学的真实感和多样性。
英文摘要
Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchronization, weakly correlated dynamics, including eyebrow movements, eye blinks, and head motion, which are essential to photorealistic facial animation, remain difficult to model faithfully and often appear static or unnaturally repetitive. We attribute this limitation to three factors: (a) insufficient conditioning for weakly correlated dynamics; (b) the limited ability of deterministic regression to capture diverse motion patterns; (c) data bottlenecks from unreliable upper-face pseudo-labels and limited dataset diversity. To address these issues, we propose SubtleTalk, a framework for generating natural and controllable weakly correlated facial dynamics via multi-condition modeling and residual flow matching. First, to compensate for the limited guidance of speech alone, we introduce interpretable controls, including prosody, regional intensity, and Valence-Arousal signals, to explicitly capture the timing, magnitude, and affective variation of weakly correlated dynamics. Second, to overcome the limited expressiveness of deterministic regression, we build residual flow matching based on a stable speech-driven motion prior, allowing the model to capture stochastic deviations beyond deterministic prediction. Third, to alleviate the data bottleneck, we construct SubtleTalk-Face, a large-scale 3D facial animation dataset comprising about 3,900 identities and 74 hours of data, built via a simple and scalable pseudo-labeling pipeline and featuring improved upper-face tracking and frame-level VA annotations. Extensive experiments demonstrate that our method significantly improves the realism and diversity of weakly correlated facial dynamics while preserving accurate lip synchronization.