发表机构
Agence Nationale de la Recherche; COMPROMIS project(法国国家科研机构; COMPROMIS项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音深度伪造检测中理解分类原因的挑战,引入音素级分析框架,利用梯度加权类激活映射结合语音识别生成显著性图,该方法在ASVspoof 5实验中检测性能可比且能提供跨条件语言解释。
AI 中文摘要
随着使用诸如wav2vec 2.0和HuBERT等自监督表示,语音深度伪造检测的准确性有所提高,但理解语音被分类为真实或深度伪造的原因仍是一个挑战。为追求更可靠和可解释的人工智能,我们引入了一个将模型预测与可测量语音单元相连接的音素级分析框架。我们的事后可解释性方法利用梯度加权类激活映射结合语音识别来生成与音素和停顿对齐的显著性图,可应用于多种基于卷积神经网络的语音深度伪造检测系统,揭示了与欺骗性语音相关的、人类可理解的、具有统计显著性的攻击和说话者相关语音线索。使用ASVspoof 5进行的实验显示,与类似架构具有可比的检测性能,同时能跨说话者和欺骗条件提供语言解释。
英文摘要
As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy and interpretable artificial intelligence, we introduce a phoneme-level analysis framework that connects model predictions to phonetic units. Our post-hoc explainability method is generally applicable to a variety of speech deepfake detection systems based on convolutional neural networks since it leverages Gradient-weighted Class Activation Mapping in conjunction with speech recognition to generate saliency maps aligned with phonemes and pauses. This pipeline reveals statistically significant attack- and speaker-dependent phonetic cues associated with spoofed speech in terms that humans can understand. Experiments using ASVspoof 5 show comparable detection performance to similar architectures while providing linguistic interpretations across speakers and spoofing conditions.
CommentsAccepted at IEEE Spoken Language Technology 2026