arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的语音克隆系统实则是一款隐形的语音匿名器

Your Voice Cloning System is Secretly a Voice Anonymizer

Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu

arXiv 2608.27360首次发表:更新:

发表机构

ZHAW School of Engineering; Centre for Artificial Intelligence(苏黎世应用科技大学工程学院; 人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将多语言语音克隆模型XTTSv2重新用于说话人匿名化,通过迭代优化策略平衡隐私与实用性,在7种欧洲语言上实现接近最优的隐私性及更优的语音质量,无需重新训练。

AI 中文摘要

说话人匿名化旨在抑制语音中可识别说话人的属性,同时保留语言内容与音质。我们提出将XTTSv2(一款在27000小时语音数据上训练的多语言语音克隆模型)重新用于说话人匿名化,无需重新训练。我们的核心见解是,XTTSv2的语音克隆能力可独立于说话人身份保留韵律结构,通过以伪说话人为条件实现语音转换。我们引入迭代优化策略,通过最大化说话人不相似度与可懂度的调和均值,平衡隐私性与实用性。在CommonVoice和Multilingual LibriSpeech数据集涵盖的7种欧洲语言上评估显示,我们的系统实现了接近最优的隐私性(EER≈0.49)、具有竞争力的可懂度,且语音质量显著优于专用匿名化基线,同时无需针对特定语言训练。我们在此处发布代码:this https URL。

英文摘要

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER $\approx$ 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑