AI 中文总结
研究针对TTS系统中稳健性、延迟和韵律的权衡问题,提出基于稀疏时间嵌入策略的StellarTTS框架及语义感知编解码器,实现对音素多方面精细控制,其轻量级模型实时因子达0.08,实验证明该框架在多方面表现出色。
AI 中文摘要
稳健性、延迟和韵律之间的权衡对文本转语音(TTS)系统提出了严峻挑战。自回归模型虽保真但速度慢且易出错;非自回归(NAR)模型虽快,但常因严格对齐牺牲韵律自然度。本文介绍了基于稀疏时间嵌入策略的新型移动优化NAR TTS框架StellarTTS,能对音素持续时间、发音和韵律进行精细控制。还提出语义感知编解码器以实现高效单阶段解码。基于稀疏时间嵌入的83M参数轻量级掩码生成变压器实时因子达0.08。实验表明,StellarTTS与现有TTS系统相比,延迟更低、稳健性更强,在音频质量、韵律自然度和说话人相似度方面也保持竞争力。
英文摘要
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.
CommentsAccepted by ASRU 2025