LibriSpeech中内容泄露及其对说话者匿名化隐私评估的影响
Content Leakage in LibriSpeech and Its Impact on the Privacy Evaluation of Speaker Anonymization
- DFKI(德意志联邦研究院)
- Technical University of Berlin(柏林技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示了LibriSpeech数据集中因说话者词汇独特性导致的身份泄露问题,并比较了EdAcc数据集在隐私保护方面的改进。
AI中文摘要:
说话者匿名化旨在隐藏说话者的身份,而无需考虑语言内容。在本研究中,我们揭示了常用于评估匿名化方法的LibriSpeech数据集的一个弱点:LibriSpeech说话者所读的书籍如此独特,以至于可以通过其词汇量来识别说话者。即使完美的匿名化方法也无法防止这种身份泄露。EdAcc数据集在这方面表现更好:只有少数说话者可以通过其词汇量被识别,促使攻击者寻找其他途径来确定匿名化说话者身份。EdAcc还包含自发性语音和更多样化的说话者,补充了LibriSpeech并提供了更多关于匿名化方法如何工作的见解。
英文摘要:
Speaker anonymization aims to conceal a speaker's identity, without considering the linguistic content. In this study, we reveal a weakness of Librispeech, the dataset that is commonly used to evaluate anonymizers: the books read by the Librispeech speakers are so distinct, that speakers can be identified by their vocabularies. Even perfect anonymizers cannot prevent this identity leakage. The EdAcc dataset is better in this regard: only a few speakers can be identified through their vocabularies, encouraging the attacker to look elsewhere for the identities of the anonymized speakers. EdAcc also comprises spontaneous speech and more diverse speakers, complementing Librispeech and giving more insights into how anonymizers work.