arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15110cs.CVcs.AI

CETalk:面向音频驱动的3D说话头生成的连续效价-唤醒度控制

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

  • Hefei University of Technology(合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li

AI总结:

本文提出CETalk框架,以连续VA表示实现音频驱动3D说话头的细粒度情感控制,构建3D-VA-MEAD数据集,实验显示其在唇形同步和情感表现力上优于现有方法,可实现平滑可控的情感过渡。

AI中文摘要:

情感3D说话头生成旨在合成具有准确唇形同步的富有表现力的面部动画,但现有方法往往依赖离散情感类别,无法捕捉情感的连续演变,还忽略了语音发音与情感表达之间的时间频率不匹配问题。本文提出CETalk,这是一个以连续效价-唤醒度(Valence-Arousal, VA)表示为条件、用于细粒度情感控制的音频驱动3D面部动画框架。CETalk通过三个关键组件预测FLAME参数序列:动态情感调制模块,利用音频衍生线索自适应缩放情感强度;多尺度时间建模机制,采用并行分支将高频发音动作与低频情感动态解耦;动态融合机制,通过自适应门控网络整合这些多尺度特征。为支持训练与评估,本文构建了3D-VA-MEAD大规模数据集,该数据集带有自动估算的VA标注及重构的3D面部动作。大量实验表明,CETalk在唇形同步准确率和情感表现力上均优于现有最优方法,同时能实现平滑且可控的情感过渡。

英文摘要:

Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

补充信息

↑