发表机构
University of Michigan-Flint(密歇根大学弗林特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示神经编解码器压缩使现有音频深度伪造检测器性能大幅下降,提出PCL-NET方法,通过成对一致性学习微调XLS-R,将平均EER从28.67%降至12.77%,并构建了Audio Neural Codec-Spoof数据集。
AI 中文摘要
现有的音频深度伪造检测(ADD)数据集和检测器主要针对基于声码器的合成构建,并针对传统的事后扰动(如MP3/AAC压缩或加性噪声)进行评估,这些扰动独立于生成过程。然而,最近的语音合成器,特别是基于ALM的系统,使用神经音频编解码器进行压缩,并作为从生成的令牌重建波形的再合成组件,产生不同于事后压缩的伪影。神经编解码器因此扮演双重角色:一些设计用于低带宽通信下的纯压缩,而另一些则用作再合成组件。尽管存在这种双重角色,对基于编解码器的压缩的鲁棒性(不同于事后压缩)在很大程度上仍未探索。我们揭示了这一差距,表明最先进的(SOTA)ADD模型在编解码器压缩的语音上性能急剧下降;特别是,使用编解码器再合成数据作为基于编解码器生成的代理进行训练的系统最为脆弱,合法压缩的真实语音经常被误分类为伪造。为了调查这一点,我们构建了Audio Neural Codec-Spoof数据集,将七种神经编解码器算法应用于现有的ADD基准,隔离编解码器诱导的再合成伪影作为基于编解码器生成的受控代理。作为基线缓解措施,我们提出了PCL-NET(成对一致性学习网络),微调预训练的XLS-R(300M)编码器,使用成对一致性目标,最小化话语未压缩和编解码器压缩版本之间的表示距离,将编解码器伪影与真实与伪造的决策分离。结果,PCL-NET在神经编解码器压缩下的平均EER从28.67%降低到12.77%,同时保持具有竞争力的CoSG-based深度伪造检测性能。我们还将接受后公开数据集在Hugging Face上。
英文摘要
Existing audio deepfake detection (ADD) datasets and detectors are primarily built for vocoder-based synthesis, evaluated against traditional post-hoc perturbations such as MP3/AAC compression or additive noise, applied independently of generation. However, recent speech synthesizers, particularly ALM-based systems, use neural audio codecs both for compression and as the resynthesis reconstructing waveforms from generated tokens, producing artifacts distinct from post-hoc compression. Neural codecs thus play a dual role: some are designed for pure compression under low-bandwidth communication, while others serve as resynthesis components. Despite this dual role, robustness to codec-based compression, unlike post-hoc compression, remains largely unexplored. We expose this gap, showing that state-of-the-art (SOTA) ADD models degrade drastically on codec-compressed speech; in particular, systems trained on Codec Resynthesized data as a proxy for codec-based generation prove most vulnerable, with legitimately compressed bona fide speech often misclassified as fake. To investigate this, we construct the Audio Neural Codec-Spoof dataset by applying seven neural codec algorithms to existing ADD benchmarks, isolating codec-induced resynthesis artifacts as a controlled proxy for codec-based generation. As baseline mitigation, we propose PCL-NET (Pairwise Consistency Learned Network), fine-tuning a pretrained XLS-R (300M) encoder with a pairwise consistency objective that minimizes the representation distance between an utterance's uncompressed and codec-compressed versions, disentangling codec artifacts from the real-versus-fake decision. As a result, PCL-NET reduces average EER under neural codec compression from 28.67% to 12.77%, while preserving competitive CoSG-based deepfake detection performance. We will also make the dataset publicly available on Hugging Face upon acceptance.