发表机构
Univ. Grenoble Alpes; CNRS; Grenoble INP(格勒诺布尔阿尔卑斯大学; 法国国家科学研究中心; 格勒诺布尔国立理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出基于内容嵌入匹配的语音匿名化模型,利用预训练编码器提取嵌入,经矢量量化和声码器解码,训练目标使匿名信号嵌入与原始匹配,引入梯度反转层丢弃说话人信息,该方法字错误率低,匿名化性能好,还能部分保留情感。
AI 中文摘要
本文提出了一种专注于保留内容而非生成逼真语音的语音匿名化模型。它依赖于从冻结的预训练wav2vec2编码器提取的内容嵌入。这些嵌入通过矢量量化和HiFi-GAN声码器解码为匿名信号,两者均在LibriTTS上训练,无任何波形重建损失或说话人嵌入映射。训练目标是使匿名信号的嵌入与原始信号的嵌入匹配。训练时,带有梯度反转层的辅助说话人分类分支用于丢弃说话人特定信息。结果表明,这种直接的基于嵌入的方法实现了非常低的字错误率(2.53),匿名化性能(等效错误率13.39)在VPC中排名第一。值得注意的是,即使没有支持训练目标,情感也能部分保留(未加权平均召回率43.91),且匿名语音可在无重建损失的情况下听到。
英文摘要
The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization and a HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddings of the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.