arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过维纳-霍普夫线性预测实现可设计解释的音频深度伪造检测

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani

arXiv 2607.12584首次发表:更新:

发表机构

Amped Software(Amped软件公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对音频深度伪造检测难题,提出基于维纳-霍普夫线性预测和轻量级二维卷积神经网络的可设计解释框架,实验显示其检测性能优、计算复杂度低,解释性分析揭示其关注要点,鲁棒性实验表明微调可应对后处理退化。

AI 中文摘要

合成语音生成方法的快速发展使音频深度伪造检测成为多媒体取证中的关键挑战。近期方法虽检测准确率高,但多依赖黑箱架构,解释性有限且计算复杂度高。本文提出基于维纳-霍普夫线性预测并经轻量级二维卷积神经网络处理的可设计解释音频深度伪造检测框架,使分类结果与信号声学特性直接透明关联。基准数据集实验表明其检测性能有竞争力且计算复杂度低得多。Grad-CAM解释性分析显示分类器关注低阶预测系数及静音和过渡区域,表明维纳-霍普夫预测器捕捉到合成语音中的混响特征和细微统计不一致。鲁棒性实验表明微调能在常见后处理退化下有效恢复检测性能。

英文摘要

The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.

CommentsAccepted at ACM IH&MMSec 2026

DOI:10.1145/3785353.3815087

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑