arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08276eess.AS

语音匿名化简化:基于投影分类器无引导的训练无关匿名化

Voice Anonymization Made Simple: Training-Free Anonymization with Projected Classifier-Free Guidance

Xiang Shi, Han Zhu, Ming Li, Xiaoxiao Miao

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种训练无关的两阶段语音匿名化方法,通过条件重构和引导抑制,在保持语言内容的同时提升说话人隐私,实验验证了其有效性。

中文摘要 AI 辅助

语音匿名化旨在隐藏说话人身份,同时保留语言和副语言信息,然而许多现有方法需要专门的训练。我们提出了一种训练无关的两阶段匿名化方法,用于预训练的流匹配语音转换系统。在条件构建阶段,源语音的内容表示在随机选择的说话人语境下被重新生成,以减少残留的源说话人信息,同时保留语言内容。在语音生成阶段,利用源语音识别与原始说话人相关的生成方向,抑制与源对齐的成分,同时保留有用的内容引导。所提出的方法仅修改推理时的条件和引导,不更新任何预训练参数。在VoicePrivacy 2026 Track 1上的实验表明,该两阶段方法在提高说话人隐私的同时,保持了具有竞争力的词错误率和情感识别性能。

英文摘要

Voice anonymization aims to conceal speaker identity while preserving linguistic and paralinguistic information, yet many existing methods require dedicated training. We propose a training-free, two-stage anonymization method for pretrained flow-matching voice conversion systems. In the condition-construction stage, the source content representation is regenerated under a randomly selected speaker context to reduce residual source-speaker information while preserving linguistic content. In the speech-generation stage, the source speech is used to identify the generation direction associated with the original speaker, and the source-aligned component is suppressed while useful content guidance is retained. The proposed method modifies only inference-time conditioning and guidance, without updating any pretrained parameters. Experiments on VoicePrivacy 2026 Track 1 show that the two-stage method improves speaker privacy while maintaining competitive word error rate and emotion recognition performance.

发表机构

  • Duke Kunshan University(杜克昆山大学)
  • Xiaomi Corporation(小米公司)
  • School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑