arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27230eess.AS

通过在 WavLM 空间中学习 kNN 传输实现一步式语音转换

One-Step Voice Conversion by Learning kNN Transport in WavLM Space

Anton Selitskiy, David Millard

首次发表
浏览论文内容

中文总结 AI 辅助

提出 kNN-FM-VC,一种基于流匹配的单一网络,在 WavLM 空间中学习 kNN 传输,实现一步式高质量语音转换,显著降低 WER。

中文摘要 AI 辅助

语音转换(VC)系统分为两大类:一类是非参数嵌入空间方法,无需训练模型,但在目标话语较短时性能下降;另一类是基于频谱图的神经架构,通过包含数千万参数的多模块流水线实现高质量转换。我们提出 kNN-FM-VC,一种单一的条件流匹配网络,学习近似源说话人与目标说话人 WavLM 嵌入分布之间的 kNN-VC 映射,用基于 kNN 生成配对训练的神经回归器替代显式的逐点 kNN 匹配。该模型通过交叉注意力和 FiLM 以目标说话人为条件,并在三种高斯条件路径(薛定谔桥、直线和恒定方差高斯管)下训练,支持少步采样。与使用上采样阶段后接 kNN 匹配的 Phoneme Hallucinator 不同,我们的 13M 参数模型通过单一学习网络执行转换,并支持一步推理。在 LibriSpeech 上,一步式高斯桥在 WER 和估计语音质量方面均优于 FreeVC 和 Phoneme Hallucinator。相对于 kNN 和 kDOT,它显著降低了 WER。

英文摘要

Voice conversion (VC) systems fall into two families: non-parametric embedding-space methods, which need no trained model but degrade on short target utterances, and spectrogram-based neural architectures, which achieve strong quality via multi-module pipelines with tens of millions of parameters. We propose kNN-FM-VC, a single conditional flow-matching network that learns to approximate the kNN-VC mapping between WavLM embedding distributions of source and target speakers, replacing explicit pointwise kNN matching with a neural regressor trained on kNN-generated pairs. The model is conditioned on the target speaker via cross-attention and FiLM, and trained under three Gaussian conditional paths (Schrödinger bridge, straight line, and constant-variance Gaussian tube), enabling few-step sampling. Unlike Phoneme Hallucinator, which uses an upsampling stage followed by kNN matching, our 13M-parameter model performs conversion with a single learned network and supports one-step inference. On LibriSpeech, the one-step Gaussian Bridge achieves lower WER and higher estimated speech quality than FreeVC and Phoneme Hallucinator. Relative to kNN and kDOT, it substantially reduces WER.

发表机构

  • Stony Brook University(石溪大学)
  • University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑