arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

告别等错误率,迎接局部信息泄露:评估针对1对N关联威胁的语音匿名化

Goodbye Equal Error Rate, Hello Local Information Disclosure: Evaluating Voice Anonymisation against 1-to-N Linkage Threats

Dāvis Šterns, Konstantinos Drossos, Natasha Fernandes, Tom Bäckström, Catuscia Palamidessi

arXiv 2607.06259首次发表:更新:

AI 中文总结

研究针对1对N关联威胁的语音匿名化,提出模块化信息理论评估框架,核心指标LID量化隐私损失,评估VoicePrivacy 2024挑战赛系统发现EER近乎完美的系统仍有局部漏洞,采用局部隐私指标很关键。

AI 中文摘要

语音匿名化旨在保护说话者身份。目前,其经验性隐私评估严重依赖等错误率(EER)。EER最初用于生物特征验证,全局聚合分数,假设攻击者仅验证两个特定语音样本是否匹配(1对1比较),这与现实世界数据库关联攻击的威胁模型不匹配。近期1对N指标解决了聚合问题,但忽略了生物特征证据的大小。本文提出专为1对N关联威胁模型设计的模块化信息理论评估框架。核心指标局部信息泄露(LID)通过校准原始相似度分数量化单次试验话语的确切隐私损失。对VoicePrivacy 2024挑战赛中表现最佳的系统评估发现,EER近乎完美的系统仍有局部漏洞,最坏情况下每次试验话语泄露达1比特。采用局部隐私指标对捕捉最坏情况风险和符合严格隐私法规至关重要。

英文摘要

Voice anonymisation aims to protect speaker identity. Currently, its empirical privacy evaluation heavily relies on the Equal Error Rate (EER). Originally designed for biometric verification, EER aggregates scores globally, implicitly assuming an attacker is only trying to verify if two specific voice samples match (a 1-to-1 comparison). This introduces a threat model mismatch with real-world database linkage attacks, where an attacker searches across a fixed set of N enrolled identities (a 1-to-N closed-set search), allowing global averages to obscure localised privacy failures. While recent 1-to-N metrics address this aggregation issue, they abstract away the magnitude of the biometric evidence. In this paper, we propose a modular, information-theoretic evaluation framework explicitly designed for the 1-to-N linkage threat model. Within this framework, our core metric, Local Information Disclosure (LID), provides a principled estimate of the identity information disclosed by a single trial utterance in bits by mapping raw similarity scores to an estimated posterior distribution over the enrolled identities. Demonstrating this framework on VoicePrivacy 2024 Challenge systems reveals how global metrics can obscure privacy vulnerabilities. Even for top-performing systems where standard evaluations report near-perfect EERs (48%), our metric exposes that attackers gain a statistical advantage in at least 63% of trials with maximum observed disclosures reaching 1 bit (effectively doubling the attacker's confidence in the target speaker after observing a single trial utterance). We conclude that shifting toward explainable metrics and appropriate threat models is a practical step toward identifying worst-case vulnerabilities and aligning with strict privacy regulations.

CommentsAccepted for publication at the Symposium on Security and Privacy in Speech Communication (SPSC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑