探究音频深度伪造检测器中的说话者身份敏感性
Probing Speaker Identity Sensitivity in Audio Deepfake Detectors
浏览论文内容
中文总结 AI 辅助
研究音频深度伪造检测器对说话者身份的敏感性问题,提出身份敏感性分数(ISS),该方法无需真实标签,通过计算检测器分数变化量化身份敏感程度,能有效预测错误分类,为基于说话者的故障分析提供实用诊断。
中文摘要 AI 辅助
音频深度伪造检测器旨在区分真实语音和合成语音,在标准基准测试中通常表现良好。然而,在一个数据集上错误率不到1%的检测器,在另一个数据集上评估时错误率可能会增加二十倍。我们认为一个因素是对说话者身份的依赖:标准训练语料库将说话者身份与真实/合成标签相关联,使检测器部分依赖与说话者相关的线索而非仅合成特征。我们提出身份敏感性分数(ISS),一种逐话语诊断方法,量化检测器输出在不同说话者身份背景下的变化程度。ISS在推理时无需真实标签,可根据检测器分数和参考说话者示例池计算得出。在两个检测器和两个数据集上,错误分类话语的ISS分数比正确分类话语高29至52倍,ISS单独预测错误分类的曲线下面积(AUC)高达0.954。为测试ISS是否真正捕捉身份敏感行为而非仅作为预测置信度的代理,我们对500个话语进行语音转换并测量检测器分数变化。被ISS标记为身份敏感的话语对这种操作的响应比标记为稳定的话语强烈19至30倍。这些结果表明ISS是音频深度伪造检测中基于说话者的故障分析的实用推理时诊断方法。
英文摘要
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing factor is speaker-identity reliance: standard training corpora correlate speaker identity with the genuine/synthetic label, allowing detectors to partially rely on speaker-related cues rather than synthesis artifacts alone. We propose the Identity Sensitivity Score (ISS), a per-utterance diagnostic that quantifies how much a detector's output changes across different speaker identity contexts. ISS requires no ground-truth labels at inference time and can be computed from the detector score and a pool of reference speaker examples. Across two detectors and two datasets, incorrectly classified utterances have ISS scores 29 to 52 times higher than correctly classified utterances, and ISS alone predicts misclassification with area-under-curve (AUC) up to 0.954. To test whether ISS actually captures identity-sensitive behavior rather than serving only as a proxy for prediction confidence, we apply voice conversion to 500 utterances and measure the resulting detector-score shift. Utterances flagged as identity-sensitive by ISS respond 19 to 30 times more strongly to this manipulation than utterances flagged as stable. These results position ISS as a practical inference-time diagnostic for speaker-dependent failure analysis in audio deepfake detection.
发表机构
- Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。