arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29021eess.AS

超越语音:用于统一全类型音频深度伪造检测的双域自监督学习融合方法

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

Cunhang Fan, Junqin Cao, Tian Gao, Zhipeng Xie, Jun Xue, Zhao Lv, Xin Fang

首次发表
浏览论文内容

中文总结 AI 辅助

针对未知音频类型的全类型音频深度伪造检测难题,本文提出双域SSL融合方法,结合EAT-large与wav2vec 2.0 XLS-R-300M特征,在AT-ADD挑战赛赛道2获95.58%宏F1值且排名第二。

中文摘要 AI 辅助

统一全类型音频深度伪造检测旨在判断输入片段是真实还是伪造的,其音频类型可能为语音、环境声、歌声或音乐。现有以语音为中心或依赖特定类型的解决方案无法适用于该场景,因为测试时的音频类型未知,而所需输出仍为单一的二分类决策。为解决这些问题,本文提出一种双域自监督学习(SSL)融合方法,将异质音频映射到共享的二分类真实性空间。选用EAT-large和wav2vec 2.0 XLS-R-300M作为互补的SSL特征源,分别提供宽泛的声学与事件级表征,以及波形级、语音和语音敏感表征。分层加权融合整合来自不同Transformer深度的多级伪造痕迹,而标记级融合形成统一特征池,无需对两个SSL流之间强制帧级对齐。融合后的标记经多头注意力统计池化汇总,再由二分类多层感知机(MLP)头进行分类。在该统一核心检测器之上应用保守语音优化后,所提交系统在AT-ADD赛道2评估集上达到95.58%的宏F1值,并在挑战赛中排名第二。

英文摘要

Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the required output is still a single binary decision. To address these issues, this paper proposes a dual-domain SSL fusion method that maps heterogeneous audio into a shared binary authenticity space. EAT-large and wav2vec 2.0 XLS-R-300M are used as complementary SSL feature sources, providing broad acoustic and event-level representations as well as waveform-level, vocal, and speech-sensitive representations. Layer-wise weighted fusion integrates multi-level artifacts from different transformer depths, while token-level fusion forms a unified feature pool without enforcing frame-level alignment between the two SSL streams. The fused tokens are summarized by multi-head attentive statistics pooling and classified with a binary MLP head. With conservative speech refinement applied on top of this unified core detector, the submitted system achieves 95.58% Macro-F1 on the AT-ADD Track 2 evaluation set and ranks second in the challenge.

发表机构

  • Anhui University(安徽大学)
  • Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta(长三角地区安全人工智能安徽省实验室)
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑