arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DriftTTS:无需蒸馏的少步文本到语音合成,基于分布匹配漂移

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam

arXiv 2610.03390首次发表:更新:

发表机构

University of Massachusetts Amherst; Worcester Polytechnic Institute(马萨诸塞大学阿默斯特分校; 伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DriftTTS,一种无需蒸馏和对抗训练的少步文本到语音合成模型,通过分布匹配漂移目标在梅尔域特征空间训练,在LJSpeech上达到与Matcha-TTS相当的性能。

AI 中文摘要

少步神经文本到语音合成模型通常依赖于缩短的扩散或流匹配调度,或依赖于从预训练的多步教师模型中蒸馏。为了避免这些依赖,我们提出了DriftTTS,一种无需生成式教师、蒸馏或对抗性判别训练的少步梅尔频谱图生成器。DriftTTS在由原始梅尔频谱和在同一LJSpeech训练分割上预训练的冻结掩码自编码器编码器定义的梅尔域特征空间中使用分布匹配漂移目标。在线策略回放训练解码器在其自身的中间状态上,并支持推理深度达到训练的回放深度。在LJSpeech上,DriftTTS在NFE=4时实现了3.87 dB的MCD和3.7%的WER,而Matcha-TTS为3.85 dB和3.4%。在完全配对的盲听测试中,DriftTTS获得了4.18的MOS,而Matcha-TTS为3.96,真实音频为4.22。这些结果证明了在没有预训练生成式教师的情况下,少步合成具有竞争力。代码可在https://github.com/BASHLab/driftTTS.git获取。

英文摘要

Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑