基于3D高斯泼溅的面部关键点引导的说话头生成
Talking Head Synthesis with Facial Landmark Guidance via 3D Gaussian Splatting
浏览论文内容
中文总结 AI 辅助
针对音频驱动说话头中嘴部运动不准确和表情细节弱的问题,提出基于3D高斯泼溅的面部关键点引导空间增强与全局补偿方法,提升视觉质量、真实感和唇同步。
中文摘要 AI 辅助
音频驱动的数字人生成在虚拟通信、沉浸式交互和媒体制作中扮演着重要角色。随着神经辐射场(NeRF)和3D高斯泼溅(3DGS)的发展,最近的说话头系统获得了更逼真的3D面部几何和外观建模。一个遗留的困难是,语音特征主要描述时间声学模式,而非明确的面部布局。因此,直接用音频驱动3D面部变形可能会产生不准确的嘴部运动、微弱的表情细节和局部伪影。为解决这一问题,我们提出了一种面部关键点引导的空间增强模块。预测的关键点为选择并丰富表情敏感面部区域周围的空间点提供了结构线索。我们进一步引入了一种全局关键点补偿机制,其中全部关键点集合被编码成一个条件向量以细化3DGS属性。这种补偿为底层形状表示提供了全脸结构信息。在自驱动和交叉驱动设置下的实验表明,所提出的方法提高了视觉质量、面部真实感和嘴唇同步性。
英文摘要
Audio-driven digital human generation plays an important role in virtual communication, immersive interaction, and media production. With the development of Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), recent talking-head systems have obtained more faithful 3D facial geometry and appearance modeling. A remaining difficulty is that speech features mainly describe temporal acoustic patterns rather than explicit facial layouts. As a result, directly driving 3D facial deformation with audio may produce inaccurate mouth motion, weak expression details, and local artifacts. To address this issue, we propose a facial-keypoint-guided spatial enhancement module. The predicted landmarks provide structural cues for selecting and enriching spatial points around expression-sensitive facial regions. We further introduce a global landmark compensation mechanism, where the full set of keypoints is encoded into a conditioning vector to refine 3DGS attributes. This compensation supplies whole-face structural information to the underlying shape representation. Experiments under self-driven and cross-driven settings show that the proposed method improves visual quality, facial realism, and lip synchronization.
发表机构
- Xinjiang University(新疆大学)
- Joint Research Laboratory for Embodied Intelligence, Xinjiang University(新疆大学具身智能联合研究实验室)
- Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(新疆大学丝绸之路多语言认知计算国际联合研究实验室)
机构由 AI 辅助整理,请以论文原文为准。