语音匿名化简化:基于投影分类器无引导的训练无关匿名化
Voice Anonymization Made Simple: Training-Free Anonymization with Projected Classifier-Free Guidance
浏览论文内容
中文总结 AI 辅助
提出一种训练无关的两阶段语音匿名化方法,通过条件重构和引导抑制,在保持语言内容的同时提升说话人隐私,实验验证了其有效性。
中文摘要 AI 辅助
语音匿名化旨在隐藏说话人身份,同时保留语言和副语言信息,然而许多现有方法需要专门的训练。我们提出了一种训练无关的两阶段匿名化方法,用于预训练的流匹配语音转换系统。在条件构建阶段,源语音的内容表示在随机选择的说话人语境下被重新生成,以减少残留的源说话人信息,同时保留语言内容。在语音生成阶段,利用源语音识别与原始说话人相关的生成方向,抑制与源对齐的成分,同时保留有用的内容引导。所提出的方法仅修改推理时的条件和引导,不更新任何预训练参数。在VoicePrivacy 2026 Track 1上的实验表明,该两阶段方法在提高说话人隐私的同时,保持了具有竞争力的词错误率和情感识别性能。
英文摘要
Voice anonymization aims to conceal speaker identity while preserving linguistic and paralinguistic information, yet many existing methods require dedicated training. We propose a training-free, two-stage anonymization method for pretrained flow-matching voice conversion systems. In the condition-construction stage, the source content representation is regenerated under a randomly selected speaker context to reduce residual source-speaker information while preserving linguistic content. In the speech-generation stage, the source speech is used to identify the generation direction associated with the original speaker, and the source-aligned component is suppressed while useful content guidance is retained. The proposed method modifies only inference-time conditioning and guidance, without updating any pretrained parameters. Experiments on VoicePrivacy 2026 Track 1 show that the two-stage method improves speaker privacy while maintaining competitive word error rate and emotion recognition performance.
发表机构
- Duke Kunshan University(杜克昆山大学)
- Xiaomi Corporation(小米公司)
- School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。