arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

任务感知的联合剪枝与蒸馏用于高效音频深度伪造检测

Task-Aware Joint Pruning and Distillation for Efficient Audio Deepfake Detection

Miao He, Peng Cheng, Zhongjie Ba, Qing Wen, Li Lu, Xin Yang, Kui Ren

arXiv 2610.05264首次发表:更新:

发表机构

Zhejiang University; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security; Shanghai Institute for Advanced Study, Zhejiang University; Qingdao Institute of Software College of Computer Science and Technology, China University of Petroleum (East China)(浙江大学; 杭州高新区(滨江)区块链与数据安全研究院; 浙江大学上海高等研究院; 中国石油大学(华东)计算机科学与技术学院青岛软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对深度伪造检测模型参数过多难以部署的问题,提出任务感知的联合剪枝与蒸馏框架,结合跨域知识蒸馏与运动引导结构化剪枝,将模型压缩至31.9M参数,FLOPs降低6.3倍,性能仅下降1.30%。

AI 中文摘要

语音合成技术的进步使得深度伪造语音越来越逼真,对安全构成日益严重的威胁。虽然基于自监督学习(SSL)的检测器达到了最先进的性能,但其计算需求(通常超过3亿参数)阻碍了在资源受限设备上的部署。现有的压缩方法主要针对以内容为中心的任务设计,在直接应用于深度伪造检测时难以保持有竞争力的性能。我们提出了一种任务感知的联合剪枝与蒸馏框架,将跨领域知识蒸馏与运动引导的结构化剪枝相结合,以迁移伪造判别性知识并在激进压缩下保留关键结构。我们的框架将模型缩减至3190万参数,计算量降低6.3倍,与未压缩的基线相比,在多个数据集上的平均性能下降仅为1.30%,展示了在设备端部署的巨大潜力。

英文摘要

Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parameters) prevent deployment on resource-constrained devices. Existing compression methods, designed mainly for content-centric tasks, struggle to maintain competitive performance when directly adapted to deepfake detection. We propose a Task-Aware Joint Pruning and Distillation framework that combines cross-domain knowledge distillation with movement-guided structured pruning to transfer forgery-discriminative knowledge and preserve critical structures under aggressive compression. Our framework reduces the model to 31.9M parameters with 6.3$\times$ FLOPs reduction, with an average performance drop of only 1.30\% across multiple datasets compared to the uncompressed baseline, demonstrating strong potential for on-device deployment.

Comments6 pages, 4 figures, accepted to Interspeech 2026

DOI:10.21437/Interspeech.2026-1766

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑