arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13842cs.SDcs.AIeess.AS

CRAF:用于深度伪造语音检测的跨视图残差感知融合

CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki

首次发表
浏览论文内容

中文总结 AI 辅助

CRAF框架利用听觉大语言模型引导的跨视图注意力与残差融合,提升深度伪造语音检测对未见攻击的泛化能力,在ASVspoof 5上取得5.96%的EER。

中文摘要 AI 辅助

语音合成和语音转换的最新进展使得深度伪造语音越来越逼真,使得对未见过的欺骗攻击的泛化成为一个关键挑战。预训练的语音和音频模型为提高对此类未见攻击的鲁棒性提供了一个有前景的方向。自监督学习(SSL)模型捕获细粒度、低层次的声学特征,而听觉大语言模型(ALLMs)提供更高层次的上下文表示。这些互补的视图可以为提高对未见攻击的泛化能力提供有用的线索。然而,直接融合并未明确地将两个视图共享的信息与视图特定的互补信息分离,限制了有效的跨视图整合。为解决这一问题,我们提出了CRAF,一种跨视图残差感知融合框架,该框架使用ALLM引导的跨视图注意力来丰富SSL表示,并采用ALLM作为高层次参考,将ALLM可解释的信息与互补的SSL残差信息分离。残差通过自适应门控进行选择性细化,并通过SSL主融合进行整合。在ASVspoof 5上的实验表明,使用Kimi-Audio的CRAF实现了5.96%的等错误率(EER)和0.1192的最小检测代价函数(minDCF),展示了其对未见欺骗攻击的鲁棒性。

英文摘要

Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge. Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks. Self-supervised learning (SSL) models capture fine-grained, low-level acoustic characteristics, whereas Auditory Large Language Models (ALLMs) provide higher-level contextual representations. These complementary views can provide useful cues for improving generalization to unseen attacks. However, direct fusion does not explicitly disentangle information shared across the two views from view-specific complementary information, limiting effective cross-view integration. To address this, we propose CRAF, a cross-view residual-aware fusion framework that uses ALLM-guided cross-view attention to enrich SSL representations and adopts ALLM as a high-level reference to separate ALLM-explainable information from complementary SSL residual information. The residual is selectively refined through adaptive gating and integrated through SSL-primary fusion. Experiments on ASVspoof 5 show that CRAF with Kimi-Audio achieves an EER of 5.96% and a minDCF of 0.1192, demonstrating robustness to unseen spoofing attacks.

发表机构

  • Japan Advanced Institute of Science and Technology(日本先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑