arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35411cs.SDcs.AI

GLAD:用于鲁棒语音深度伪造检测的全局-局部自适应检测器

GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu

AI总结:

提出GLAD全局-局部自适应检测器,通过分层全局-局部骨干网络捕获局部伪造痕迹,引入分层自适应门控机制应对层重要性变化,并采用SaniBoost数据增强策略,显著提升语音深度伪造检测在未见领域的鲁棒性。

AI中文摘要:

近年来,基于人工智能的语音合成技术取得了显著进展,能够生成高度逼真的语音,这使得语音深度伪造检测(SDD)在防止滥用方面的重要性日益增加。虽然主流的基于自监督学习(SSL)的检测器取得了强劲的性能,但它们对未见领域的泛化能力较差,并且由于偏向全局语义一致性,常常忽略细粒度的信号伪影。在本文中,我们首次进行了详细的实证和可视化分析,以明确验证这些局限性。我们的调查揭示了两个关键的架构脆弱性:(1)系统性地无法捕获局部伪造痕迹,以及(2)严重缺乏对SSL层重要性领域驱动变化的适应性,导致静态聚合策略容易过拟合。为了解决这些脆弱性,我们提出了全局-局部自适应检测器(GLAD)。具体而言,为了捕获局部伪造,GLAD采用了一种分层全局-局部(HGL)骨干网络,通过融合全局语言和声学特征与细粒度的局部信号细节,明确弥合了粒度差距。为了应对分布外(OOD)场景中的层重要性变化,我们引入了一种分层自适应门控(HAG)机制,该机制以样本特定的方式动态重新校准层级关注。最后,为了解决环境偏差引起的捷径学习问题,我们引入了SaniBoost,一种复合数据增强策略,用于稳健的信号标准化和噪声净化。大量实验表明,GLAD显著优于最先进的方法,尤其是在未见领域上。代码将在发表后发布。

英文摘要:

Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.

↑