发表机构
Communication University of China; Ant Group; Machine Intelligence, Ant Group; Institute of Automation, Chinese Academy of Sciences; Beijing Institute of Technology(中国传媒大学; 蚂蚁集团; 蚂蚁集团机器智能; 中国科学院自动化研究所; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文总结了ACM MM 2026 AT-ADD全类型音频深度伪造检测挑战赛的两个赛道任务、设计、结果及参赛系统的常见模式,指出当前仍存在泛化、鲁棒性及性能均衡等挑战。
AI 中文摘要
本文总结了ACM多媒体2026 AT-ADD全类型音频深度伪造检测挑战赛。AT-ADD包含两个赛道:一是在真实声学和信道变化下的鲁棒语音深度伪造检测,二是对语音、环境声、歌声及音乐的类型无关检测。我们描述了挑战赛任务、数据集与评估集设计、官方排行榜结果,以及参赛系统中观察到的常见设计模式。Track 1的最佳系统在最终评估集上达到90.71%的Macro-F1,Track 2的最佳系统达到96.10%的Macro-F1。最终提交结果显示,强系统通常结合大规模自监督音频表征、数据增强、多裁剪推理及结构化融合或路由。结果还揭示了在泛化到未见过的生成器、对真实语音域失真的鲁棒性,以及跨异构音频类型的均衡性能方面仍存在的挑战。
英文摘要
This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard results, and common design patterns observed in participating systems. The best Track 1 system achieved 90.71% Macro-F1 on the final evaluation set, while the best Track 2 system achieved 96.10% Macro-F1. The final submissions show that strong systems commonly combine large-scale self-supervised audio representations, data augmentation, multi-crop inference, and structured fusion or routing. The results also reveal remaining challenges in generalization to unseen generators, robustness to realistic speech-domain distortions, and balanced performance across heterogeneous audio types.
CommentsAccepted to ACM MM 2026