发表机构
Soonchunhyang University(顺天乡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建多主干自监督集成模型完成ImageCLEF 2026音频深度伪造检测任务并夺冠生成子任务,分析发现生成-检测不对称性,指出部署的主要挑战是11.25%的误报差距。
AI 中文摘要
本文介绍了“Go-To-Germany”团队参与ImageCLEF 2026音频深度伪造检测与生成任务的情况。我们的检测系统基于四主干自监督学习(SSL)集成模型构建,融合了WavLM-Large、Wav2Vec2-XLS-R-300M、ECAPA-TDNN和x-vector表示,在ImageCLEF 2026官方评估中取得了0.9522的最终得分,对参与者生成的深度伪造样本准确率达1.0000,对保留的主办方真实数据准确率为0.8875。对于生成子任务,我们的官方团队提交结果是经均匀混响处理的F5-TTS v1基线模型,作为故意反取证探针提交,最终得分0.4304(词错误率WER为4.99%,字符错误率CER为2.07%);论文中还介绍了我们的四模型方案(GLM-TTS、F5-TTS、XTTS v2、CosyVoice3),官方提交结果即来自该方案。我们开展了跨赛道分析,揭示了显著的不对称性:我们的检测系统识别出100%的参与者生成深度伪造,而我们的官方生成提交结果虽在音频生成子任务中排名第一,且规避了61.4%的参与者检测器和56.2%的主办方检测器,但其最终得分0.4304与检测侧得分0.9522形成对比。我们还报告了基于证伪的消融实验(留一 speakers 56交叉验证、三区域主干几何、自助法置信区间及主成分分析),为多主干SSL集成的架构保险假设提供支撑。我们补充了五条跨赛道见解和五条预注册的证伪实验,将生成侧规避与检测侧设计决策关联,还公开报告了在保留的主办方真实录音上11.25%的误报差距,这是部署的主要开放挑战。
英文摘要
This paper describes the participation of team "Go-To-Germany" in the ImageCLEF 2026 Audio Deepfake Detection and Generation task. Our detection system, built on a four-backbone self-supervised learning (SSL) ensemble combining WavLM-Large, Wav2Vec2-XLS-R-300M, ECAPA-TDNN, and x-vector representations, achieved a final score of 0.9522 on the official ImageCLEF 2026 evaluation, with perfect accuracy (1.0000) on participant-generated deepfakes and 0.8875 on the held-out organizer ground-truth real data. For the Generation sub-task, our official team submission, an F5-TTS v1 baseline processed with a uniform reverberation pass and submitted as a deliberate anti-forensic probe, ranked first with a final score of 0.4304 (word error rate (WER) 4.99%, character error rate (CER) 2.07%); details of our four-model program (GLM-TTS, F5-TTS, XTTS v2, CosyVoice3), from which the official entry was drawn, appear in the paper. We present a cross-track analysis revealing a pronounced asymmetry: our detection system identifies 100% of participant-generated deepfakes, while our official generation entry, despite ranking first in the Audio Generation sub-task and evading 61.4% and 56.2% of participant and organizer detectors, attains a Final Score of 0.4304 against 0.9522 on the Detection side. We further report falsification-based ablation experiments (LOSO 56-speaker cross-validation, three-region backbone geometry, bootstrap confidence intervals, and PCA analysis) that motivate our architectural-insurance hypothesis for multi-backbone SSL ensembling. We complement these results with five cross-track insights and five pre-registered falsification experiments connecting generation-side evasion to detection-side design decisions, and we openly report an 11.25% false-positive gap on held-out organizer real recordings as the principal open challenge for deployment.
CommentsAccepted at CLEF 2026, ImageCLEF-Deepfake task. Published in CEUR-WS CLEF 2026 Working Notes