CuteTTS:基于连续隐变量自回归建模的高效高保真语音合成
CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
浏览论文内容
中文总结 AI 辅助
CuteTTS是结合语义对齐因果VAE隐变量等的连续自回归TTS系统,经引导步蒸馏后,在保持相当质量的同时降低了延迟,实现了高保真与低延迟的平衡。
中文摘要 AI 辅助
零样本文本到语音(TTS)技术现已支持交互式助手、个性化媒体及无障碍工具。所有TTS系统均需实现准确的语言渲染、一致的说话人身份及低延迟响应。然而,紧凑的流式系统必须在可预测的低速率隐变量序列中保留足够的声学细节,而迭代扩散采样与无分类器引导会在每一步自回归推理中增加推理成本。为在高保真合成与低延迟推理间取得平衡,本文提出CuteTTS,一种紧凑的连续自回归TTS系统。它结合了语义对齐的因果VAE隐变量、补丁级自回归、显式说话人条件及双向流匹配头。我们进一步引入引导步蒸馏技术,将无分类器引导与多个求解器步整合为单个区间条件学生模型。在LibriSpeech与Seed-TTS-Eval上的评估表明,该模型在零样本语音克隆中具备竞争力的可懂度与说话人相似度;与基准模型相比,蒸馏后首音频延迟降低23.3%,实时因子降低40.8%,同时保持相当的客观与主观质量。这些结果为实现兼顾高保真生成与实时交互延迟需求的连续自回归TTS提供了可行路径。
英文摘要
Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency response. Yet compact streaming systems must preserve sufficient acoustic detail in a predictable low-rate latent sequence, while iterative diffusion sampling and classifier-free guidance multiply inference cost at every autoregressive step. To strike a balance between high-fidelity synthesis and low-latency inference, we present CuteTTS, a compact continuous-autoregressive TTS system. It combines semantically aligned causal VAE latents with patch-level autoregression, explicit speaker conditioning, and a bidirectional flow-matching head. We further introduce guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student. Evaluations on LibriSpeech and Seed-TTS-Eval demonstrate competitive intelligibility and speaker similarity in zero-shot voice cloning, while distillation lowers first-audio latency by 23.3% and real-time factor by 40.8% relative to the base model with comparable objective and subjective quality. These results provide a practical path toward continuous-autoregressive TTS that reconciles high-fidelity generation with the latency demands of real-time interaction.
发表机构
- Shanghai Innovation Institute(上海创新研究院)
- OPPO AI Center(OPPO人工智能中心)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。