arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33709eess.AScs.SD

域自适应双门控专家混合用于泛化语音深度伪造检测

Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection

  • The Hong Kong Polytechnic University(香港理工大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Siqing Qin, Zhe Li, Kong Aik Lee, Man-Wai Mak

AI总结:

提出域自适应双门控专家混合框架,利用Sinc层和域原型指导路由,在跨数据集基准上实现高达40.8%的EER降低,提升语音深度伪造检测的泛化能力。

AI中文摘要:

近期语音深度伪造检测(SDD)的进展利用了专家混合(MoE)来增强泛化能力。然而,现有的门控网络往往忽略了深度伪造的声学和时间线索。在这项工作中,我们提出了一种新颖的域自适应双门控专家混合(DADGMoE)框架,用于在未见攻击类型和声学条件下进行SDD。我们创新的双门控机制利用基于Sinc层的滤波器来处理低级声学信号(原始波形)和来自大型自监督学习(SSL)模型的高级语音表征。它进一步结合了域原型,以基于隐式深度伪造模式指导专家路由。轻量级仿射专家处理路由输入。实验表明,我们的DADGMoE显著优于基线,在具有挑战性的跨数据集基准上实现了高达40.8%的相对等错误率(EER)降低。该框架展示了优越的泛化能力和高效设计。

英文摘要:

Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.

补充信息

↑