基于E-Branchformer的解耦全局-局部特征学习用于音频深度伪造检测
Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection
查看机构详情
- National University of Singapore(新加坡国立大学)
- Hanoi University of Science and Technology(河内理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出基于E-Branchformer的并行双分支架构,利用自监督语音表示,通过多头自注意力和卷积捕获全局与局部特征,集成深度卷积和压缩-激励模块增强分类令牌,在ASVspoof 2021等数据集上取得最优性能。
中文摘要 AI 辅助
语音合成技术(如文本到语音和语音转换)的快速发展对基于语音的身份验证系统构成了重大威胁,因此迫切需要鲁棒的深度伪造检测方法。在本工作中,我们提出了一种新颖的基于E-Branchformer的架构,该架构有效利用自监督语音表示进行音频深度伪造检测。我们的模型采用并行分支,通过多头自注意力同时捕获全局上下文依赖关系,并通过卷积处理捕获局部时间模式。为了增强判别能力,我们集成了深度卷积和压缩-激励模块,在特征合并后用精细的补丁令牌信息丰富分类令牌。在ASVspoof 2021 LA、DF和In-the-Wild数据集上的大量实验表明,我们的方法达到了最先进的性能,等错误率分别为0.88%、1.85%和6.30%,大幅优于现有方法。全面的消融研究验证了双分支架构提供了互补的判别信息,压缩-激励聚合显著改善了SSL特征集成,且DWConv和SE模块的组合对于有效的类令牌增强至关重要。在真实场景中的优越性能表明,该方法对多样化的声学条件和未见过的欺骗攻击具有强大的泛化能力。
英文摘要
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervised speech representations for audio deepfake detection. Our model employs parallel branches to simultaneously capture global contextual dependencies through multi-head self-attention and local temporal patterns through convolutional processing. To enhance discriminative capability, we integrate depthwise convolution and Squeeze-and-Excitation modules that enrich the classification token with refined patch token information after feature merging. Extensive experiments on ASVspoof 2021 LA, DF, and In-the-Wild datasets demonstrate state-of-the-art performance with equal error rates of 0.88%, 1.85%, and 6.30% respectively, substantially outperforming existing methods. Comprehensive ablation studies validate that the dual-branch architecture provides complementary discriminative information, Squeeze-and-Excitation Aggregation significantly improves SSL feature integration, and the combination of DWConv and SE modules is critical for effective class token enhancement. The superior performance on real-world scenarios demonstrates strong generalization capability to diverse acoustic conditions and unseen spoofing attacks.