驯服长文本语音合成
Taming Long-form Text-to-Speech
浏览论文内容
中文总结 AI 辅助
针对长文本TTS性能退化问题,提出LACI推理方法,通过近实时错误检测与回滚重生成,显著提升长文本WER和语音克隆可靠性。
中文摘要 AI 辅助
长文本语音合成(TTS)支持多轮对话,具有一致的韵律,并能从更长的参考音频中实现更高质量的语音克隆。最近的开源自回归TTS模型,如Qwen3-TTS和VoxCPM2,在短文本提示上达到了最先进的词错误率(WER)和说话人相似度(SIM),但在长文本提示下性能显著下降。我们提出了局部注意力约束推理(LACI),一种仅推理的方法,能够近实时地检测TTS错误,回滚到错误起始点并在临时护栏下重新生成,仅增加可忽略的计算开销。使用LACI,我们将Qwen3-TTS-0.6B在超过1500词提示上的10个随机种子的最差N WER从35.2%改善到3.4%,甚至超过了其在少于500词提示上的5.4%的短文本可靠性。为了展示LACI在语音克隆可靠性上的有效性,我们提出了SIM指标的滑动窗口版本,称为wSIM。wSIM揭示了SIM未捕获的几种新的失败模式。LACI将120秒参考音频上的最差N wSIM从0.01提高到0.47,同时将WER高于30%的灾难性生成率从26%降低到1%以下。
英文摘要
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4\% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1%
发表机构
- Argmax, Inc.(Argmax公司)
机构由 AI 辅助整理,请以论文原文为准。