衡量说话人去识别化系统中的软生物特征泄露
Measuring Soft Biometric Leakage in Speaker De-Identification Systems
- National Institute of Standards and Technology(国家标准与技术研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出软生物特征泄露分数(SBLS),通过直接属性推理、互信息关联检测和子群体鲁棒性三要素,量化说话人去识别化系统对零样本推理攻击的抵抗能力,并实验证明五个现有系统均存在显著软生物特征泄露漏洞。
AI中文摘要:
我们使用“再识别”一词来指从匿名化语音输出中恢复原始说话人身份的过程。说话人去识别化系统旨在降低再识别风险,但大多数评估仅关注个体层面的指标,而忽视了软生物特征泄露带来的更广泛风险。我们提出软生物特征泄露分数(SBLS),这是一种统一的方法,用于量化对非唯一特征的零样本推理攻击的抵抗能力,这些特征包括信道类型、年龄段、方言、说话人性别或说话风格。SBLS整合了三个要素:使用预训练分类器进行直接属性推理、通过互信息分析进行关联检测,以及跨交叉属性的子群体鲁棒性。通过使用公开可用的分类器应用SBLS,我们表明所有五个被评估的去识别化系统都表现出显著的脆弱性。我们的结果表明,仅使用预训练模型的对手——无需访问原始语音或系统细节——仍然能够可靠地从匿名化输出中恢复软生物特征信息,这暴露了标准分布度量无法捕捉的根本性弱点。
英文摘要:
We use the term re-identification to refer to the process of recovering the original speaker's identity from anonymized speech outputs. Speaker de-identification systems aim to reduce the risk of re-identification, but most evaluations focus only on individual-level measures and overlook broader risks from soft biometric leakage. We introduce the Soft Biometric Leakage Score (SBLS), a unified method that quantifies resistance to zero-shot inference attacks on non-unique traits such as channel type, age range, dialect, sex of the speaker, or speaking style. SBLS integrates three elements: direct attribute inference using pre-trained classifiers, linkage detection via mutual information analysis, and subgroup robustness across intersecting attributes. Applying SBLS with publicly available classifiers, we show that all five evaluated de-identification systems exhibit significant vulnerabilities. Our results indicate that adversaries using only pre-trained models - without access to original speech or system details - can still reliably recover soft biometric information from anonymized output, exposing fundamental weaknesses that standard distributional metrics fail to capture.