arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15690cs.SDcs.AIcs.LGcs.MM

为文本-音频-视频模型添加带单个零初始化层的语音克隆功能

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

  • Kandinsky Lab(坎丁斯基实验室)

机构由 AI 辅助整理,请以论文原文为准。

Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov

AI总结:

本文在基础文本-音频-视频模型上添加单个零初始化线性层,经微调后实现语音克隆,在674个说话人-文本对基准上优于5个基线,且推理时可跳过视频路径提速约30倍。

AI中文摘要:

文本到音频-视频(T2AV)生成模型可根据文本描述生成视频及其音频,但无法控制输出中的说话人是谁。本文展示,通过在其音频主干顶部添加单个零初始化线性层、进行相对较短训练时长的微调,并在推理时基于一段短参考录音,基础T2AV模型可被转化为语音克隆模型。该参考通过两种互补信号注入:其扩散潜变量被前置到音频流,且全局说话人嵌入会调制目标音频的 token。在包含30位说话人、共674个说话人-文本对的基准测试中,我们与5个强大的语音克隆文本到语音基线进行对比:我们增强后的5B模型在三个独立验证网络(ECAPA-TDNN、WavLM-SV、Resemblyzer)上均达到最高的说话人编码器余弦相似度(SECS),在统计上显著优于所有基线。该架构的附带好处是,推理时可在不使用视频路径的情况下评估音频路径,与完整的音频-视频扩散循环相比,速度提升约30倍,同时保留语音克隆行为。

英文摘要:

Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time. The reference is injected through two complementary signals: its diffusion latents are prepended to the audio stream, and a global speaker embedding modulates token of the target audio. On a benchmark of 674 speaker-text pairs spanning 30 speakers we compare against five strong voice-cloning text-to-speech baselines: our enhanced 5B model attains the highest speaker-encoder cosine similarity (SECS) across three independent verification networks (ECAPA-TDNN, WavLM-SV, Resemblyzer), statistically significantly outperforming every baseline. A side product of the architecture is that the audio path can be evaluated without the video path at inference time, yielding a ~30x speed-up over the full audio-video diffusion loop while preserving the voice-cloning behaviour.

↑