arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本到语音合成中语音表示的弗雷歇距离损失

Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu, Hung-yi Lee

arXiv 2607.06027首次发表:更新:

发表机构

Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对少步TTS模型训练问题,提出语音表示弗雷歇距离损失(SR-FD),在微调时让模型用相同采样器合成语音,将其特征与参考统计量匹配,无需判别器和推理计算,显著降低Seed-TTS英语的字错误率,是提高可懂度的分布正则化方法。

AI 中文摘要

少步扩散和流匹配文本到语音(TTS)模型通常使用局部目标进行训练,如条件流匹配、重建和停止预测。这些损失提供了稳定的优化,但未考虑采样语音是否遵循高质量语音的分布。我们提出了语音表示弗雷歇距离损失(SR-FD),用于无分词器流匹配自回归TTS的训练时分布正则化。在微调期间,模型使用与部署时相同的少步采样器合成语音,SR-FD将此语音的冻结Whisper和CTC特征的均值和协方差与从三个互补内容目标离线计算的参考统计量进行匹配。该损失无需判别器和推理时计算。在Seed-TTS英语上,四步SR-FD微调将字错误率从原始四步VoxCPM2基线的2.2279%降至1.4147%,相对降低36.5%,并超过原始十步基线的1.7366%。SR-FD是一种用于少步TTS的提高可懂度的分布正则化方法。

英文摘要

Few-step diffusion and flow-matching text-to-speech (TTS) models are usually trained with local objectives, such as conditional flow matching, reconstruction, and stop prediction. These losses provide stable optimization, but they never ask whether sampled speech follows the distribution of high-quality speech. We propose Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS. During fine-tuning, the model synthesizes speech with the same few-step sampler used at deployment, and SR-FD matches the mean and covariance of frozen Whisper and CTC features of this speech to reference statistics computed offline from three complementary content targets. The loss requires no discriminator and no inference-time computation. On Seed-TTS English, four-step SR-FD fine-tuning reduces WER from the original four-step VoxCPM2 baseline's 2.2279% to 1.4147%, a 36.5% relative reduction, and also surpasses the original ten-step baseline at 1.7366%; both gains are significant under an utterance-level paired bootstrap. Speaker similarity and objective quality proxies are preserved at the ten-step level, and an error analysis shows the gain comes from content substitutions across all prompt lengths. SR-FD is thus an intelligibility-improving distributional regularizer for few-step TTS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑