arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享状态局部变换用于免训练语音转换

Shared-State Local Translations for Training-Free Voice Conversion

Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans

arXiv 2610.01952首次发表:更新:

AI 中文总结

提出StateVC,通过共享状态局部变换实现免训练语音转换,在LibriSpeech上取得最低WER和CER,并保持高说话人相似度。

AI 中文摘要

在一次性免训练语音转换(VC)中,源话语和参考话语可能包含不同的语言内容,因此不能假设它们之间存在可靠的帧级对应关系。我们提出StateVC,它从源话语和参考话语的池化帧级WavLM表示中共同定义一组公共局部区域;我们将这些区域称为状态。这些共享状态通过将针对特定说话人对的混合高斯模型拟合到池化表示上获得,无需显式的源-参考帧匹配。在每个状态内,StateVC在原始WavLM空间中估计源到参考的均值偏移。然后,源帧的后验概率组合各状态特定的偏移,使得不同帧可以接收不同的局部更新。在LibriSpeech一次性协议下,StateVC在评估系统中实现了最低的词错误率(WER)和字符错误率(CER),分别为8.01%和3.22%,说话人相似度(SIM)为0.9512。它还在评估系统中实现了最高的平均感知说话人相似度,以及在评估的免训练系统中实现了最高的平均自然度。

英文摘要

In one-shot training-free voice conversion (VC), the source and reference utterances may contain different linguistic content, so reliable frame-level correspondence between them cannot be assumed. We propose StateVC, which jointly defines a common set of local regions from pooled frame-level WavLM representations of the source and reference utterances; we refer to these regions as states. These shared states are obtained by fitting a pair-specific Gaussian mixture model to the pooled representations, without explicit source--reference frame matching. Within each state, StateVC estimates a source-to-reference mean shift in the original WavLM space. Source-frame posterior probabilities then combine the state-specific shifts so that different frames can receive different local updates. For the LibriSpeech one-shot protocol, StateVC achieves the lowest word error rate (WER) and character error rate (CER) among the evaluated systems, at 8.01% and 3.22%, respectively, with a speaker similarity (SIM) of 0.9512. It also achieves the highest mean perceived speaker similarity among the evaluated systems and the highest mean naturalness among the evaluated training-free systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑