arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38658eess.AScs.CLcs.SD

Tacit-TTS:从自回归解码到掩码预测以实现高效的免转录语音克隆

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning

  • Dolby Laboratories(杜比实验室)

机构由 AI 辅助整理,请以论文原文为准。

Jian Chen, You Zhang, Mark Vinton

AI总结:

针对自回归TTS延迟高、非自回归依赖转录的问题,提出Tacit-TTS,用掩码预测替代自回归解码,实现免转录零样本克隆,速度提升10倍以上,且支持跨语言与非词汇参考。

AI中文摘要:

具有自回归语义建模的TTS系统展现了强大的零样本语音克隆性能和丰富的表达变化,但其顺序解码导致了显著的延迟。非自回归替代方案提供了更快的生成速度,但通常依赖于更严格的参考条件,例如在推理时需要参考语音的转录文本。我们提出了Tacit-TTS,这是一个从IndexTTS2蒸馏而来的高效免转录零样本语音克隆系统。我们的模型用掩码非自回归生成取代了自回归文本到语义解码,引入了无需训练的声学长度估计,并通过ReFlow蒸馏加速了流匹配渲染器。在两个英语和两个普通话数据集上,Tacit-TTS在生成超过5秒的语音时,实现了与IndexTTS2相当的零样本质量,同时生成速度快10倍以上。其免转录条件还支持跨语言和非词汇参考。我们使用来自其他八种语言、婴儿咿呀声和合成胡言乱语的参考验证了这一能力,在这些情况下,依赖转录的系统往往因不可靠的ASR转录而性能下降或失败。

英文摘要:

TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.

补充信息

↑