AT-ADD:面向鲁棒全类型音频深度伪造检测的基准与挑战赛
AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
浏览论文内容
中文总结 AI 辅助
本文提出AT-ADD基准与挑战赛,设双赛道分别评估鲁棒语音深度伪造检测与全类型音频深度伪造检测,官方基线及获胜系统取得对应性能,同时揭示泛化关键及未解决问题。
中文摘要 AI 辅助
近期音频生成模型可合成高保真的语音、环境音、歌声与音乐,给多媒体信任带来新风险。现有音频深度伪造检测(ADD)基准仍以语音为核心,往往未充分涵盖现实信道变化与多样音频类型。本文提出AT-ADD,这是一个大规模基准与挑战赛,旨在评估鲁棒语音深度伪造检测及全类型音频深度伪造检测。赛道1评估在未见过的生成器、多样录音条件、信号扰动与重放效应下的二分类语音检测;赛道2评估在测试时音频类型未知的情况下,对语音、声音、歌声与音乐的类型无关真假检测。我们详述数据集构建、评估协议与可复现基线,并分析提交至ACM Multimedia 2026 Grand Challenge的最终系统。最强官方基线在赛道1和赛道2评估集上的Macro-F1分别为76.73%和79.47%,而挑战赛获胜系统分别达到90.71%和96.10%。除整体排名外,对前五名提交结果的样本级分析考察了生成器与类型级难度、跨系统错误互补性及排名稳定性。结果表明,大规模自监督表征、条件感知增强、多裁剪推理及结构化融合或路由是泛化的核心,而生成器特定鲁棒性与多样音频类型间的一致性能仍是未解决的问题。
英文摘要
Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types. This paper presents AT-ADD, a large-scale benchmark and challenge designed to evaluate both robust speech deepfake detection and all-type audio deepfake detection. Track 1 evaluates binary speech detection under unseen generators, diverse recording conditions, signal perturbations, and replay effects. Track 2 evaluates type-agnostic real/fake detection over speech, sound, singing, and music when the audio type is unknown at test time. We detail the dataset construction, evaluation protocol, and reproducible baselines, and analyze the final systems submitted to the ACM Multimedia 2026 Grand Challenge. The strongest official baseline obtains 76.73% and 79.47% Macro-F1 on the Track 1 and Track 2 evaluation sets, respectively, whereas the winning challenge systems reach 90.71% and 96.10%. Beyond aggregate rankings, sample-level analysis of the top five submissions examines generator- and type-level difficulty, cross-system error complementarity, and ranking stability. The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.
发表机构
- Communication University of China(中国传媒大学)
- Ant Group(蚂蚁集团)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Beijing Institute of Technology(北京理工大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。