Tacit-TTS:从自回归解码到掩码预测以实现高效的免转录语音克隆
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对自回归TTS延迟高、非自回归依赖转录的问题,提出Tacit-TTS,用掩码预测替代自回归解码,实现免转录零样本克隆,速度提升10倍以上,且支持跨语言与非词汇参考。
AI中文摘要:
具有自回归语义建模的TTS系统展现了强大的零样本语音克隆性能和丰富的表达变化,但其顺序解码导致了显著的延迟。非自回归替代方案提供了更快的生成速度,但通常依赖于更严格的参考条件,例如在推理时需要参考语音的转录文本。我们提出了Tacit-TTS,这是一个从IndexTTS2蒸馏而来的高效免转录零样本语音克隆系统。我们的模型用掩码非自回归生成取代了自回归文本到语义解码,引入了无需训练的声学长度估计,并通过ReFlow蒸馏加速了流匹配渲染器。在两个英语和两个普通话数据集上,Tacit-TTS在生成超过5秒的语音时,实现了与IndexTTS2相当的零样本质量,同时生成速度快10倍以上。其免转录条件还支持跨语言和非词汇参考。我们使用来自其他八种语言、婴儿咿呀声和合成胡言乱语的参考验证了这一能力,在这些情况下,依赖转录的系统往往因不可靠的ASR转录而性能下降或失败。
英文摘要:
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.