发表机构
School of Software Technology, Zhejiang University; College of Computer Science and Technology, Zhejiang University; Ningbo Global Innovation Center, Zhejiang University(浙江大学软件学院; 浙江大学计算机科学与技术学院; 浙江大学宁波全球创新中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GemTalk框架结合隐式表示与显式几何先验,通过V-AEP、D-GPG和GEM模块,实现情感说话人脸的可控生成,在照片真实感与情感动态上表现优异。
AI 中文摘要
音频驱动的情感说话人脸生成旨在合成具有丰富面部动态的真实视频。然而,现有方法难以平衡可控性与视觉保真度。尽管隐式表示能捕捉丰富语义,但缺乏结构引导,常导致情感表达平淡;而显式几何方法虽能更好控制面部表情,却往往牺牲高频纹理细节。为解决该问题,我们提出GemTalk,这是一种基于扩散的框架,结合了隐式表示的语义丰富性与显式几何先验的结构精确性。我们引入视觉引导音频情感投影(V-AEP)模块提取隐式情感唇形与表情特征,同时,基于扩散的几何先验生成器(D-GPG)生成带身份感知的 blendshape 系数作为显式结构先验。关键在于,我们的几何引导情感调制(GEM)模块利用这些几何先验重新校准隐式特征的幅度,实现对情感表达(尤其是情感强度)的精确、连续控制,且不牺牲视觉质量。大量实验表明,GemTalk在照片真实感和面部情感动态方面取得了优异性能。
英文摘要
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit representations capture rich semantics, they lack structural guidance, often resulting in averaged emotional expressions. In contrast, explicit geometric methods offer better control over facial expressions but tend to sacrifice high-frequency texture details. To address it, we propose GemTalk, a diffusion-based framework that combines the semantic richness of implicit representations with the structural precision of explicit geometric priors. We introduce a Vision-guided Audio Emotion Projection (V-AEP) module to extract implicit emotional lip and expression features. At the same time, a Diffusion-based Geometric Priors Generator (D-GPG) generates identity-aware blendshape coefficients as explicit structural priors. Crucially, our Geometry-guided Emotion Modulation (GEM) module leverages these geometric priors to recalibrate the magnitude of implicit features, enabling precise, continuous control over emotional expressions, especially emotion intensity, without sacrificing visual quality. Extensive experiments show GemTalk achieves superior performance in photo-realism, and facial emotional dynamics.
Comments17 pages, 11 figures, accepted by ACMMM 2026