基于Style-Diffusion TTS模型潜在空间适配的零样本人脸到语音合成
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
浏览论文内容
中文总结 AI 辅助
该研究提出基于StyleTTS 2的人脸到语音合成框架,通过人脸适配器实现静态人脸到语音的零样本生成,在LRS3数据集上验证了效果,且映射具有语言无关性。
中文摘要 AI 辅助
零样本文本到语音(TTS)可通过简短音频提示克隆语音,但依赖参考音频的特性在仅能获取视觉信息时(如针对历史人物或游戏角色)存在障碍。本研究提出人脸到语音(F2S)框架,可从静态人脸图像生成合理语音。轻量人脸适配器结合人脸编码器上层块的软调优,将人脸识别特征与冻结的StyleTTS 2模型风格空间对齐,训练期间StyleTTS 2保持冻结状态。在大规模英语TED演讲视听语料库LRS3的保留身份上评估,合成语音自然度高(UTMOS 3.7-4.0,匹配或超过真实语音的3.61),人脸到语音检索结果始终高于随机水平,生成语音与目标说话人一致。无需重新训练,经英语训练的适配器也可生成流利西班牙语语音,表明人脸到风格的映射在很大程度上与语言无关。
英文摘要
Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Face-to-Speech (F2S) framework that predicts a plausible voice from a static facial image. A lightweight Face Adapter, together with soft-tuning of the face encoder's upper blocks, aligns face-recognition features with the style space of a frozen StyleTTS 2 model, kept frozen during training. We evaluate on held-out identities from LRS3, a large-scale audiovisual corpus of English TED-talk videos. The synthesized speech is highly natural (UTMOS 3.7-4.0, matching or exceeding the 3.61 of ground truth), face-to-voice retrieval is consistently above chance, and the generated voice is consistent with the target speaker. Without any retraining, an English-trained adapter also produces fluent Spanish speech, indicating that the face-to-style mapping is largely language-agnostic.
发表机构
- Universitat Oberta de Catalunya (UOC)(开放加泰罗尼亚大学(UOC))
- Monoceros Labs(麒麟座实验室)
- University of Granada(格拉纳达大学)
机构由 AI 辅助整理,请以论文原文为准。