arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05507eess.AS

AffectDF:针对情感表达型攻击的最全面语音深度伪造检测基准

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

Aurosweta Mahapatra, Xiutian Zhao, Shreeram Suresh Chandra, Zihan Zhang, Zongyang Du, Ismail Rasim Ulgen, Kong Aik Lee, Nicholas Andrews, Carlos Busso, Berrak Sisman

AI总结:

该研究提出涵盖多类情感表达型语音伪造攻击的AffectDF基准,实验发现常规训练的SDD系统在该基准上鲁棒性严重下降,为开发更鲁棒的检测模型提供支撑。

AI中文摘要:

语音深度伪造检测(SDD)系统在常规基准上表现出色,但现有数据集对情感表达型攻击及近期基于大音频语言模型(LALM)的攻击覆盖有限。现有情感欺骗数据集在规模和攻击多样性上也存在局限,通常仅涵盖语音转换(VC)或文本到语音(TTS)攻击。本文提出AffectDF——针对情感表达型语音深度伪造的最全面基准,涵盖TTS、VC、情感VC及基于LALM的欺骗攻击,涉及表演性和自发性情感语音。AffectDF包含约260小时语音,由5种情感状态下的21种欺骗攻击生成。我们在常规和情感欺骗条件下对最先进的SDD系统进行基准测试,包括对仅推理提示和监督微调的基于LALM的检测器的评估。实验显示,在常规基准上训练的模型在AffectDF上评估时会出现严重的鲁棒性下降,多个系统性能接近随机水平。令人惊讶的是,即使大规模情感训练也未能持续提升跨域鲁棒性,表明当前SDD系统无法在情感和韵律变异性下学习到通用的欺骗表示。鲁棒性还在情感状态、攻击家族及表演性与自发性情感语音条件间存在显著差异。这些发现暴露了当前SDD系统的根本局限,并确立AffectDF为开发更鲁棒欺骗检测模型的基准。

英文摘要:

Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.

↑