arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Not Quite My Tempo:面向唇形同步配音的语音活动感知语音合成

Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova, Milana Shkhanukova, Andreas Kellner, Mattia Antonino Di Gangi

arXiv 2609.26486首次发表:更新:

发表机构

AppTek GmbH(AppTek有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出以二值语音活动信号替代唇部运动作为条件,用于唇形同步配音的语音合成,实现高精度时序跟随与自然韵律,并支持推理时可选约束。

AI 中文摘要

自动唇形同步配音要求语音合成模型生成与源片段时序精确匹配的目标语言交替语音和静音模式,以确保最佳观看体验。先前的工作通过将语音合成过程条件于从视频信号中提取的唇部运动来解决此问题。在本工作中,我们将语音生成条件于一个二值语音活动信号,该信号具有轻量级表示,并可通过多种方式产生。我们通过广泛的主观和客观评估表明,模型能以高精度跟随语音活动信号,同时保持自然韵律和句内语义恰当的停顿位置。通过在训练期间随机掩蔽该条件,我们使该特征在推理时完全可选,允许编辑在需要时强制执行或放宽唇形同步约束。

英文摘要

Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.

Commentsaccepted at Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑