可学习的无分类器引导空嵌入用于增强可控语音合成
Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis
浏览论文内容
中文总结 AI 辅助
本文提出用可学习的无条件嵌入替代固定空向量进行无分类器引导,提升语音合成的可控性,并在说话人相似度、稳定性和表现力上优于固定方法。
中文摘要 AI 辅助
无分类器引导(CFG)在文本到语音(TTS)系统中被广泛采用,通过在有条件预测和无条件预测之间进行插值来增强生成质量和条件保真度。一种常见的无条件技术是使用空表示,即固定的空向量。在这项工作中,我们提出用可学习的无条件嵌入替换这种表示,该嵌入被优化为表示有意义的无条件状态。客观和主观评估表明,可学习的空嵌入在说话人相似度、语音稳定性和表现力方面始终优于固定的空嵌入,同时对更大的引导尺度表现出更强的鲁棒性。我们进一步表明,为每个TTS条件模态学习一个不同的无条件嵌入,可以对说话人和文本引导进行细粒度控制,展示了生成语音中相似度与质量、稳定性与表现力之间的权衡。
英文摘要
Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and unconditioned predictions. A common unconditional technique is to use an empty representation, in the form of a fixed null vector. In this work, we propose replacing this representation with a learnable unconditional embedding, optimized to represent a meaningful unconditional state. Objective and subjective evaluations demonstrate that learnable null embeddings consistently outperform fixed null embeddings across speaker similarity, speech stability, and expressiveness, while exhibiting greater robustness to larger guidance scales. We further show that learning a distinct unconditional embedding for each of the TTS conditioning modalities allows fine-grained control over speaker and text guidance, showcasing the trade-off between similarity and quality, and stability and expressiveness in the generated speech.
发表机构
- Cantina Labs(坎蒂纳实验室)
机构由 AI 辅助整理,请以论文原文为准。