发表机构
Soongsil University; Seoul National University(崇实大学; 首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频到语音合成中的一对多映射问题,提出文本条件框架WYS,利用注意力融合模块整合文本与视频,结合条件流匹配,在LRS2/3上达到音视频同步最优并保持高文本准确率。
AI 中文摘要
视频到语音合成旨在从无声的说话人脸视频中生成自然流畅且语音学上准确的语音。该任务的一个基本挑战是固有的“一对多”映射问题,即视觉动态往往缺乏足够的信息来唯一确定对应的语音内容。为解决此问题,我们提出了 Watch Your Speech (WYS),一种将文本条件作为显式语言线索以缓解视觉模糊性的视频到语音合成框架。我们的框架包含一个基于注意力的嵌入融合模块,该模块将文本上下文与视频序列协同整合,并配合条件流匹配目标以实现高保真语音生成。在 LRS2 和 LRS3 数据集上的大量实验表明,WYS 取得了优越的性能,在音视频同步(LSE-C/D)方面创下了新的最先进结果,同时保持了极具竞争力的文本准确性(WER)。主观评估进一步证实,我们的模型生成的语音具有接近人类的自然度,验证了文本条件在内容可控的视频到语音合成中的有效性。项目页面:此 https URL
英文摘要
Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech
CommentsAccepted to BMVC 2026