arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StellarTTS:用于低延迟和鲁棒语音合成的稀疏时间嵌入

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He, Yujie Hou, Xiaoyang Xing, Huiyan Li, Bing Han, Yanmin Qian

arXiv 2607.19859首次发表:更新:

AI 中文总结

研究针对TTS系统中稳健性、延迟和韵律的权衡问题,提出基于稀疏时间嵌入策略的StellarTTS框架及语义感知编解码器,实现对音素多方面精细控制,其轻量级模型实时因子达0.08,实验证明该框架在多方面表现出色。

AI 中文摘要

稳健性、延迟和韵律之间的权衡对文本转语音(TTS)系统提出了严峻挑战。自回归模型虽保真但速度慢且易出错;非自回归(NAR)模型虽快,但常因严格对齐牺牲韵律自然度。本文介绍了基于稀疏时间嵌入策略的新型移动优化NAR TTS框架StellarTTS,能对音素持续时间、发音和韵律进行精细控制。还提出语义感知编解码器以实现高效单阶段解码。基于稀疏时间嵌入的83M参数轻量级掩码生成变压器实时因子达0.08。实验表明,StellarTTS与现有TTS系统相比,延迟更低、稳健性更强,在音频质量、韵律自然度和说话人相似度方面也保持竞争力。

英文摘要

The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.

CommentsAccepted by ASRU 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑