arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03436cs.GRcs.CVcs.MM

语音的形态:面向语音驱动三维面部动画的协同发音几何度量

The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

Danzel Serrano, Przemyslaw Musialski

首次发表
浏览论文内容

中文总结 AI 辅助

针对语音驱动三维面部动画中协同发音被简化的问题,提出基于唇部路径长度的几何度量,验证了现有方法轨迹平坦性缺陷,并指明改进方向。

中文摘要 AI 辅助

语音驱动的三维面部动画能够再现可识别的口型。然而,它可能简化口型之间的运动,而该运动承载着协同发音,即每个音素周围的音素如何塑造其发音方式。我们引入了一种针对这种轨迹塑形的几何度量:唇部路径长度与语音片段中元音、辅音和元音位置间最短路径的比值。与端点弦长相比,这种考虑辅音的路径既考虑了必经的过渡,又避免了退化,同时保持了对均匀运动增益的不变性。该度量仅需强制对齐,因此适用于不存在真值(ground truth)的场景。我们在四种最先进的方法上进行了演示,每种方法对应一种架构家族,涵盖实时和离线方法。所有四种方法生成的唇部轨迹都比真实采集的语音更平坦。在与帧率匹配的真值对比中,DiffPoseTalk、ARTalk和FaceFormer表现出明显的缺陷,在此度量下相当于移除了真实语音快速发音成分的15%至60%。CodeTalker在主要度量上表现边缘,而在辅助度量上表现明显。一项预先注册的研究(97名观众,3523个判断)支撑了所测方向:对真实运动进行受控阻尼会降低得分并受到惩罚,而夸张在测试范围内未检测到惩罚。观众在73.4%的句子比较中偏好真实语音,并且在单词层面总体上也是如此。综合来看,该度量、其校准以及该研究识别出了感知相关的轨迹塑形损失,并为改进合成发音提供了具体目标。

英文摘要

Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment. In contrast to the endpoint chord, this consonant-aware route accounts for obligatory transit and avoids degeneracy, while preserving invariance to uniform motion gain. The measure needs only a forced alignment, so it applies where no ground truth exists. We demonstrate it on four state-of-the-art methods, one per architectural family, real-time and offline. All four trace flatter lip trajectories than captured speech. Against frame-rate-matched ground truth, DiffPoseTalk, ARTalk and FaceFormer show clear deficits, equivalent on this measure to removing 15-60% of real speech's fast articulatory component. CodeTalker is marginal on the primary measure and clear on a companion measure. A pre-registered study with 97 viewers and 3,523 judgments underpins the measured direction: controlled damping of real motion lowers the score and is penalized, whereas exaggeration shows no detected penalty over the tested range. Viewers also prefer real speech in 73.4% of sentence comparisons and, in the aggregate, on single words. Together, the measure, its calibration and the study identify a perceptually relevant loss of trajectory shaping and a concrete target for improving synthesized articulation.

发表机构

  • New Jersey Institute of Technology(新泽西理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑