发表机构
University of Bucharest; Technical University Munich; University of Tübingen(布加勒斯特大学; 慕尼黑工业大学; 蒂宾根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视听深度伪造检测的泛化难题,提出DF-MoE框架,通过多预训练模型提取多模态线索并结合MoE主干网络,在五个基准数据集上实现优于现有方法的检测性能。
AI 中文摘要
视听深度伪造检测是当前的研究热点,其核心挑战之一是开发能跨不同深度伪造生成方法泛化的检测器。我们推测,通过预训练模型从可用的音频和视觉模态中提取多个高层线索,可缓解过拟合问题。因此,我们组装了多种预训练模型,用于提取嘴部动作、面部解析、面部表情、头部姿态、视线跟踪、心率、音频情感及语音活动等特征。我们进一步通过混合专家(Mixture-of-Experts,MoE)主干网络整合单模态和多模态线索,以检测深度伪造。我们在五个深度伪造检测基准数据集(MAVOS-DD、AVLips、PolyGlotFake、BioDeepAV、FakeAVCeleb)上开展了域内和跨域实验,将我们的框架DF-MoE与当前最优方法进行对比。结果表明,DF-MoE取得了更优的深度伪造检测结果,优于所有对比方法。我们的代码已在该httpsURL开源。
英文摘要
Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion and speech activity. We further integrate both unimodal and multimodal cues via a Mixture-of-Experts (MoE) backbone to detect deepfakes. We perform in-domain and cross-domain experiments on five benchmarks for deepfake detection (MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV, FakeAVCeleb) to compare our framework (DF-MoE) with state-of-the-art methods. Our results indicate that DF-MoE obtains superior deepfake detection results, surpassing all competing methods. We release our code at https://github.com/vladhondru25/DF-MoE.
CommentsAccepted at BMVC 2026