用于稳健音频深度伪造检测的神经音频编解码器
Neural Audio Codec for Robust Audio Deepfake Detection
- Yonsei University(延世大学)
- MAAP Lab(MAAP实验室)
- The University of Tokyo(东京大学)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对低比特率编码导致音频深度伪造检测性能下降的问题,提出保留取证信息的神经音频编解码器FP-NAC,通过检测器引导微调在ASVspoof 2019 LA上将EER最多降低49.8个百分点,并提升多个检测器性能。
AI中文摘要:
音频深度伪造检测器通常在未压缩音频上进行评估,尽管现实世界中的音频经常经过低比特率编码。在本工作中,我们研究了音频编码如何跨编解码器、比特率和检测器影响深度伪造检测,发现在较低比特率下错误率更高。一种混合配对协议隔离了编码引起的真实和伪造音频的变化,揭示了不对称的、依赖于编解码器的失败模式:低比特率下的DAC和EnCodec主要降低真实音频检测性能,而X-Codec则表现出更强的伪造侧限制。受这些发现启发,我们提出了一种保留取证信息的神经音频编解码器(FP-NAC),它使用检测器引导的目标微调预训练编解码器,同时保留其原生硬量化路径和比特率。在ASVspoof 2019 LA上,与原始DAC在0.5 kbps下相比,FP-NAC将等错误率(EER)降低了高达49.8个百分点,同时保持了可比的重建质量。尽管仅由一个检测器监督,FP-NAC在多个检测器上均提升了性能,突显了取证透明度作为编解码器设计目标与感知质量并重的重要性。我们的代码可在该https URL获取。
英文摘要:
Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8~pp compared with the original DAC at 0.5~kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.