超越EER:说话人去标识化中信息泄漏的多维评估
Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
浏览论文内容
中文总结 AI 辅助
本文提出一个包含五个互补指标的多维评估框架,用于全面衡量说话人去标识化系统中的信息泄漏,并证明单一指标(如EER)不足以准确反映隐私保护效果。
中文摘要 AI 辅助
说话人去标识化(SDID)旨在通过隐藏说话人身份来保护隐私,同时保持语音的可用性。然而,当前的评估往往将隐私简化为单一维度——生物特征验证性能,通常以等错误率(EER)来衡量。这种狭隘的关注忽略了关键的泄漏渠道,如软生物特征推断、嵌入级重新识别和结构模板相似性,这些渠道威胁到生物特征参考的不可关联性和不可逆性。我们提出了一个涵盖五个互补指标的全面评估框架:(i)EER,(ii)软生物特征泄漏分数,(iii)累积匹配特征重新识别分析,(iv)典型相关分析和Procrustes嵌入对齐,以及(v)通过词错误率和语义相似性衡量的可理解性。通过评估来自IARPA ARTS项目的五个SDID系统,我们证明这些指标捕捉了信息泄漏的独立维度。我们的结果表明,依赖单一指标可能错误地描述SDID系统的隐私属性。
英文摘要
Speaker de-identification (SDID) aims to preserve privacy by concealing speaker identity while maintaining speech utility. However, current evaluations often reduce privacy to a single dimension - biometric verification performance - typically measured by Equal Error Rate (EER). This narrow focus ignores critical leakage channels, such as soft biometric inference, embedding-level re-identification, and structural template similarity, which threaten the unlinkability and irreversibility of biometric references. We propose a holistic evaluation framework across five complementary metrics: (i) EER, (ii) soft biometric leakage score , (iii) cumulative match characteristic re-identification analysis, (iv) canonical correlation analysis and Procrustes embedding alignment, and (v) intelligibility via word error rate and semantic similarity. Evaluating five SDID systems from the IARPA ARTS program, we demonstrate that these metrics capture independent dimensions of information leakage. Our results indicate that reliance on a single metric can misrepresent the privacy properties of an SDID system.
发表机构
- National Institute of Standards and Technology(美国国家标准与技术研究院)
- Chakra Consulting Inc.(Chakra咨询公司)
机构由 AI 辅助整理,请以论文原文为准。