arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更快的IndexTTS-2:在GPU上加速和流式自回归零样本文本到语音合成

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

Muyang Du, Shuang Yu, Junjie Lai

arXiv 2607.21042首次发表:更新:

AI 中文总结

研究旨在加速自回归文本到语音模型IndexTTS-2,利用NVIDIA TensorRT和TensorRT-LLM在GPU上实现生产部署,支持流式合成与批处理推理,实验表明在Seed-TTS基准测试中能显著加速且质量损失小,为加速类似模型提供参考。

AI 中文摘要

自回归文本到语音模型自然度高,但由于顺序生成令牌导致推理速度慢,限制了其在低延迟生产应用中的部署。IndexTTS-2是一种先进的自回归TTS模型,由GPT、流匹配扩散变压器和声码器组成。尽管合成质量高,但在没有流或批处理支持的情况下,其推理速度几乎无法达到实时。我们提出了更快的IndexTTS-2,它使用NVIDIA TensorRT和TensorRT-LLM加速IndexTTS-2的所有神经网络组件,以便在GPU上进行生产部署。更快的IndexTTS-2还支持对延迟敏感的交互式应用程序进行流式合成,并对所有组件进行批处理推理,以最大限度地提高GPU利用率。在英文和中文的Seed-TTS基准测试中进行的实验表明,自回归GPT的加速比高达5.0倍,端到端加速比为3.6倍,同时在单词错误率、说话者相似度和自然度方面的下降最小。我们的方法为在GPU上有效加速类似的自回归语音模型提供了实用参考。

英文摘要

Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2 is a state-of-the-art autoregressive TTS model consisting of a GPT, a flow-matching Diffusion Transformer, and a vocoder. Despite its high synthesis quality, its inference speed barely reaches real-time without streaming or batching support. We present Faster IndexTTS-2, which accelerates all neural network components of IndexTTS-2 for production deployment on GPUs using NVIDIA TensorRT and TensorRT-LLM. Faster IndexTTS-2 also enables streaming synthesis for latency-sensitive interactive applications, and batched inference across all components to maximize GPU utilization. Experiments on the Seed-TTS benchmark for both English and Chinese demonstrate up to 5.0$\times$ speedup on the autoregressive GPT and 3.6$\times$ end-to-end, with minimal degradation in word error rate, speaker similarity, and naturalness. Our methodology provides a practical reference for efficiently accelerating similar autoregressive speech models on GPUs.

Comments4 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑