arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语音合成内容表征的综合研究

A Comprehensive Study of Content Representations for Speech Synthesis

Diego Torres, Axel Roebel, Nicolas Obin

arXiv 2609.30975首次发表:更新:

发表机构

STMS Lab; IRCAM; CNRS; Sorbonne Université(STMS实验室; 声学/音乐研究和协作学院; 法国国家科学研究中心; 索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过仅以各内容表征为条件的生成模型,在通用框架下比较SSL特征、监督式token等表征,发现解耦取决于训练目标与信息容量的交互,监督式表征仅在容量受限时才能解耦说话人身份。

AI 中文摘要

语音内容表征是语音转换、语音到语音翻译和多模态语言模型的核心,然而它们很少在一个直接衡量每种表征所包含内容的通用生成框架下进行比较。我们通过训练一个仅以每种表征为条件的生成模型,并沿着内容、说话人身份和韵律三个维度评估生成的音频来解决这一问题。在SSL特征、监督式token、后验图和神经音频编解码器中,我们发现了两种截然不同的模式:几乎能重建原始音频的表征,以及能有效解耦说话人身份的表征。这些结果表明,解耦不仅取决于监督本身,还取决于训练目标与表征信息容量之间的相互作用:监督式表征只有在容量受到充分约束时才能解耦说话人身份。

英文摘要

Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.

Comments5 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑