arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于语音和环境声音深度伪造检测的组件级集成融合

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection

André Runewicz, Karla Schäfer, Martin Steinebach

arXiv 2607.16369首次发表:更新:

AI 中文总结

针对ICME 2026 ESDD2挑战赛的音频五类分类任务,提出基于四个预训练模型的组件级集成系统,经微调、训练增强变体及融合选定检查点等操作,取得较好成绩,在31个团队中排第5,超越官方基线。

AI 中文摘要

本文介绍了我们提交给ICME 2026 ESDD2环境感知语音和声音深度伪造检测挑战赛的内容。该任务要求对音频片段进行五类分类,其中语音、环境声音、两者或两者都不被伪造。我们提出了一个基于四个公开可用的预训练反伪造模型的组件级集成系统:XLSR-Mamba、DF-Arena、SLS和TCM-ADD。每个模型在官方CompSpoofV2开发数据上进行微调,使用三个二元头进行原始、语音和环境声音检测。我们进一步训练RawBoost增强变体,并使用边际空间分数融合组合选定的检查点。具有轻量级头部和类别偏差校准的组件级融合策略产生了我们的最佳配置,在评估集上达到0.7715的宏F1,在测试集上达到0.7828的宏F1,在最终排名阶段在31个团队中排名第5,大幅超过官方基线。

英文摘要

This paper describes our submission to the ICME 2026 ESDD2 challenge on environment-aware speech and sound deepfake detection. The task requires five-class classification of audio clips in which speech, environmental sound, both components, or neither component may be spoofed. We propose a component-level ensemble system based on four publicly available pre-trained anti-spoofing models: XLSR-Mamba, DF-Arena, SLS, and TCM-ADD. Each model is fine-tuned on the official CompSpoofV2 development data using three binary heads for original, speech, and environmental sound detection. We further train RawBoost-augmented variants and combine selected checkpoints using margin-space score fusion. A component-wise fusion strategy with lightweight head- and class-bias calibration yields our best configuration, reaching 0.7715 macro-F1 on the evaluation set and 0.7828 macro-F1 on the test set, ranking 5th out of 31 teams in the final ranking phase and substantially outperforming the official baseline.

CommentsAccepted to 2026 ICME workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑