arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SphereVAE:用于稳健自回归语音表示建模的超球面潜变量自编码器

SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling

Haoyu Zhang, Jingbin Hu, Hanke Xie, Qirui Zhan, Wenhao Li, Ziyu Zhang, Xiaming Ren, Yue Li, Xunyu Zhu, Zhipeng Chen, Lei Xie

arXiv 2609.09903首次发表:更新:

发表机构

Northwestern Polytechnical University; NetEase Inc.(西北工业大学; 网易公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SphereVAE通过将VAE潜空间约束到单位超球面,为自回归语音生成提供有界预测目标,缓解误差累积与潜变量漂移,提升长文本生成稳定性。

AI 中文摘要

随着语音生成技术的快速发展,离散编解码器表示因其提供了稳定的预测范式而被广泛使用。然而,在表现性语音生成中,离散编解码器的量化瓶颈导致了细粒度韵律、音色、发音和帧间连续性方面的信息缺口。连续表示(如VAE潜变量)通过消除这一约束,已成为自回归建模中更有效的替代方案。但当连续表示被用作自回归预测目标时,预测误差会沿着生成链累积,导致潜变量漂移并降低长文本生成的稳定性。为缓解这一问题,我们提出了SphereVAE,它将VAE潜空间约束在单位超球面上。SphereVAE在超球面上定义了幂球面后验分布,并将潜变量分布正则化向均匀先验,使得信息主要通过方向变化来编码,为自回归预测提供了有界的几何目标,降低了范数漂移的风险。由于潜变量自由度降低,SphereVAE在重建指标上不如标准VAE。然而,当集成到VoxCPM中进行零样本语音合成和长文本生成时,它在保持相似说话人相似度的同时,取得了更低的内容错误率,并展现出更稳定的长程说话人一致性。这些结果表明,适当的潜变量几何约束能有效缓解语音生成中自回归误差累积和漂移问题。

英文摘要

With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronunciation, and frame-to-frame continuity. Continuous representations (e.g., VAE latents), by eliminating this constraint, have emerged as a more effective alternative for autoregressive modeling. Yet when continuous representations are used as autoregressive prediction targets, prediction errors can accumulate along the generation chain, causing latent drift and degrading long-form stability. To mitigate this problem, we propose SphereVAE, which constrains the VAE latent space to the unit hypersphere. SphereVAE defines a Power Spherical posterior on the hypersphere and regularizes the latent distribution toward a uniform prior, so that information is encoded mainly by directional variation, providing a bounded geometric target for autoregressive prediction and reducing the risk of norm drift. SphereVAE underperforms the standard VAE on reconstruction metrics due to reduced latent freedom. However, when integrated into VoxCPM for zero-shot TTS and long-text generation, it yields lower content error rates with comparable speaker similarity, and shows more stable long-range speaker consistency. These results indicate that an appropriate latent geometric constraint can effectively mitigate autoregressive error accumulation and drift in speech generation.

Comments15 pages, 4 figures. Accepted to NCMMSC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑