X2Streaming-TTS:基于流式文本的因果令牌级文本语音合成(含语音状态继承)
X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
浏览论文内容
中文总结 AI 辅助
本文提出X2Streaming-TTS因果TTS框架,通过因果承诺与语音状态继承实现严格令牌级流式语音合成,性能优于伪流式模型,TTFT表现优异且质量接近离线基线。
中文摘要 AI 辅助
流式文本语音合成是低延迟口语对话系统的核心需求,但许多系统需等待句子级文本,仅为伪流式。真正的令牌级合成需从不确定前缀生成语音,同时在无界流中保持感知连续性,且上下文有界。本文提出X2Streaming-TTS,一款因果TTS框架,可处理异步到达的文本令牌并生成语音,无需访问未来输入。为处理不确定前缀,引入因果承诺(causal commitment):通过感知不确定性的缓冲机制暂存模糊表达,执行容量自适应、标点感知的分段;为保持声学连续性,进一步引入因果语音状态继承(causal speech-state inheritance):跨段边界传递完整的Code2Wav状态及选定的历史说话人(Talker)状态。结合注意力先验约束,该框架既阻止访问未来位置,又保留有界声学上下文。实验显示,X2Streaming-TTS在多数主观与客观指标上优于现有伪流式模型;进一步分析表明,因果承诺可稳定在线分段,减少因上下文不足导致的失败,语音状态继承则可提升边界连续性,且不降低自然度或说话人身份。X2Streaming-TTS实现了严格的令牌级合成,质量与评估的离线基线相当,单请求的首次音频令牌时间(TTFT)中位数为15.8 ms,128个并发请求时的TTFT中位数为260.8 ms,其实现已公开。
英文摘要
Streaming text-to-speech is essential for low-latency spoken dialogue systems, yet many systems wait for sentence-level text and are therefore only pseudo-streaming. True token-level synthesis must generate speech from uncertain prefixes while maintaining perceptual continuity over an unbounded stream with bounded context. We present X2Streaming-TTS, a causal TTS framework that consumes asynchronously arriving text tokens and emits speech without accessing future input. To handle uncertain prefixes, we introduce causal commitment, which keeps ambiguous expressions provisional through uncertainty-aware buffering and performs capacity-adaptive, punctuation-aware segmentation. To preserve acoustic continuity, we further introduce causal speech-state inheritance, which carries the complete Code2Wav state and selected historical Talker states across segment boundaries. Together with an attention prior constraint, it blocks access to future positions while retaining bounded acoustic context. Experiments show that X2Streaming-TTS outperforms existing pseudo-streaming models on most subjective and objective metrics. Further analysis shows that causal commitment stabilizes online segmentation and reduces failures caused by insufficient context, while speech-state inheritance improves boundary continuity without degrading naturalness or speaker identity. X2Streaming-TTS thus achieves strict token-level synthesis with quality comparable to the evaluated offline baselines, a median time to first audio token (TTFT) of 15.8 ms for a single request, and a median TTFT of 260.8 ms at 128 concurrent requests. Our implementation is publicly available at https://github.com/X-Square-Robot/X2Streaming-TTS .
发表机构
- X Square Robot(X Square机器人公司)
机构由 AI 辅助整理,请以论文原文为准。