arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SemBridge:用于连续隐变量自回归语音生成的语义令牌锚定

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

Hanke Xie, Haopeng Lin, Jiale Qian, Dake Guo, Yuepeng Jiang, Zhichao Wang, Wenxiao Cao, Jingbin Hu, Guobin Ma, Wenhao Li, Huakang Chen, Chengyou Wang, Ming Tao, Zhonghua Fu, Lei Xie, Xinsheng Wang

arXiv 2608.07462首次发表:更新:

发表机构

Soul AI Lab(Soul AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SemBridge框架,通过离散语义令牌监督自回归LM状态,在零样本TTS和SVS任务中提升内容准确性,同时保持说话人相似度与感知质量,为连续语音生成提供有效方向。

AI 中文摘要

连续隐变量自回归语音生成已成为离散令牌建模的有前景替代方案,它避免了量化损失并保留了更丰富的声学信息。然而,连续声学目标并未将语言结构作为显式令牌级预测目标呈现,因此自回归语言模型(LM)必须通过声学预测间接获取语言结构,这可能损害生成语音的内容保真度。我们提出SemBridge,一种用于连续隐变量自回归语音生成的仅训练用语义令牌锚定框架。SemBridge使用离散语义令牌直接监督自回归LM状态,并采用语义对齐声学VAE在相同语义参考下组织连续目标空间。语义监督仅在训练期间使用,推理阶段完全保持连续。我们在零样本文本到语音(TTS)和评分条件歌声合成(SVS)上评估SemBridge,在多个基准测试中,SemBridge提高了内容准确性(以词错误率和字符错误率WER/CER衡量),同时保持了有竞争力的说话人相似度和感知质量。实验结果表明,为自回归状态学习提供显式语义令牌监督是连续语音生成的有效且通用方向。语音样本可用,模型代码和检查点将在此httpsURL lab/SemBridge提供。

英文摘要

Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit token-level prediction tar- gets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acous- tic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation. SemBridge uses discrete se- mantic tokens to directly supervise autoregressive LM states and employs a Semantic-Aligned Acoustic VAE to organize the continuous target space under the same semantic refer- ence. The semantic supervision is used only during train- ing, while inference remains entirely continuous. We evalu- ate SemBridge on zero-shot text-to-speech (TTS) and score- conditioned singing voice synthesis (SVS). Across multi- ple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and percep- tual quality. Experimental results demonstrate that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation. Speech samples are available.1 The model code and checkpoints will be available at https://github.com/ASLP- lab/SemBridge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑