发表机构
Southeast University; School of Computer Science and Engineering(东南大学; 计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3DGS说话头渲染的漏嘴伪影问题,提出PD-GS模型,通过LFM融合音频与音素信息,在HDTF数据集上实现最优嘴唇几何,提升神经化身的语言忠实度。
AI 中文摘要
3D高斯溅射(3DGS)可实现快速、逼真的说话头渲染,但准确的嘴唇关节运动仍难以实现:嘴部动作常被过度平滑,且可能违反双唇闭合等硬性发音约束,产生臭名昭著的“漏嘴”伪影。一个关键难点在于,短暂、离散的发音事件是在回归目标下从连续声学嵌入中推断出来的,这会使预测偏向平均的嘴部配置。尽管现代自监督语音编码器提供了丰富的韵律和语音线索,但它们并未提供明确的、帧对齐的语言目标,无法可靠地消除闭合级事件的歧义。我们提出音素驱动高斯溅射(PD-GS),它通过自动ASR和强制对齐管道获得的时间对齐音素令牌来增强3DGS说话头。我们的核心组件是语言融合模块(LFM),它通过学习到的门控机制自适应地将连续音频上下文与离散音素嵌入融合,使模型既能保留音频驱动的平滑动态,又能在发音关键段加强音素引导。PD-GS仅使用单目视频通过图像重建和嘴唇地标监督进行训练。在HDTF数据集上,PD-GS在对比基线中实现了最佳的嘴唇几何(LMD为2.66),并在具有挑战性的音素序列中定性减少了闭合违规,产生了更符合语言忠实度的神经化身。
英文摘要
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbf{Phoneme-Driven Gaussian Splatting (PD-GS)}, which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbf{Linguistic Fusion Module (LFM)}, adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.
CommentsAccepted to ACM MM 2026