arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MADBench:面向模态感知音频深度伪造检测的基准

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang

arXiv 2608.09593首次发表:更新:

AI 中文总结

MADBench是首个区分语音与环境音频的音频深度伪造检测基准,测试发现环境音频操纵更易检测,现有预训练检测器在两类声学成分上均失效,为相关研究提供严谨基础。

AI 中文摘要

近期语音合成与音频生成技术的进步,使得高保真声学伪造的成本降低且难以溯源,催生了一种现实的攻击场景:在原本真实的视频中,语音和背景音频可被独立操纵。然而现有研究要么聚焦于视觉操纵,要么孤立处理语音检测,要么将语音与非语音音频混为单一未区分的音频流,忽略了背景音频带来的独特取证挑战。这种混同具有重要影响:这两种声学成分源于根本不同的生成机制,呈现出不同的伪影特征,对检测系统构成不同挑战。我们推出MADBench,首个将语音与环境音频视为不同声学成分的基准,支持对跨独立操纵伪造源的音频深度伪造检测进行成分感知评估。我们在统一协议下对代表性的最先进检测器和多模态大语言模型进行基准测试。实验结果表明,在通用编码器下,环境音频操纵比合成语音更易检测;现有预训练检测器在两种声学成分上均失效;被操纵的环境音频会不对称地降低语音深度伪造检测性能,这些发现在以往基准的单标签范式下完全无法被观察到。MADBench为未来研究稳健的、成分感知的音频深度伪造检测奠定了严谨基础。

英文摘要

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.

Comments11 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑