arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

证据子空间投影:衡量自我监督语音模型中多少证据可解释深度伪造检测

Evidence Subspace Projection: Measuring How Much Evidence Explains Deepfake Detection in Self-Supervised Speech Models

Yixuan Xiao, Cheng-Wei Lin, Xin Wang, Yassine El Kheir, Arnab Das, Tim Polzehl, Sebastian Möller, Ngoc Thang Vu

arXiv 2607.11538首次发表:更新:

发表机构

University of Stuttgart; National Institute of Informatics; German Research Center for Artificial Intelligence (DFKI); Technical University of Berlin(斯图加特大学; 日本信息处理研究所; 德国人工智能研究中心; 柏林技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何将自我监督学习模型用于音频深度伪造检测,提出证据子空间投影方法,通过在共享空间表示证据因素和真实性标签,投影决策向量量化证据解释力,经多数据集多设置评估,验证方法并获新见解。

AI 中文摘要

自我监督学习(SSL)模型被广泛用作最先进的音频深度伪造检测的特征提取器,但尚不清楚如何直接和定量地将SSL模型捕获的内容与检测决策联系起来。为填补这一空白,我们提出证据子空间投影方法,该方法在由SSL模型的神经元激活模式构建的共享空间中表示证据因素(如攻击类别、编解码器、性别、传输)和真实性标签。通过将决策向量投影到每个证据子空间上,我们获得一个标量比率,量化每种证据类型的解释力。我们在多个数据集上对原始、微调及后训练设置下的SSL模型进行评估。结果证实了现有研究的发现,验证了所提方法,并揭示了模型行为的新见解。

英文摘要

Self-supervised learning (SSL) models are widely used as feature extractors for state-of-the-art audio deepfake detection, but it remains unclear how to directly and quantitatively connect what SSL models capture to detection decisions. To address this gap, we propose Evidence Subspace Projection, a method that represents both evidence factors (e.g., attack category, codec, gender, transmission) and authenticity labels in a shared space constructed from SSL models' neuron activation patterns. By projecting the decision vector onto each evidence subspace, we obtain a scalar ratio that quantifies the explanatory power of each evidence type. We evaluate SSL models in raw, fine-tuned, and post-trained settings on multiple datasets. The results confirm findings from established studies, validating the proposed method, and reveal new insights into model behavior.

CommentsAccepted to Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑