发表机构
Warsaw University of Technology(华沙理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DFD-Lab模块化流水线,集成三种音视频深度伪造检测器,通过实验揭示训练增强提升AUROC但降低准确率,强调跨数据集检测挑战及指标互补性。
AI 中文摘要
比较音视频深度伪造检测器需要协调数据集适配、时间输入表示、模型接口和实验条件。我们提出了DFD-Lab,一个模块化流水线,它分离了这些职责,同时支持共享的训练和评估工作流。我们集成了三种实现:基于Xception的最大对数融合、带有时间LSTM融合的ResNet,以及我们的AVFF重实现。实验涵盖了外部测试、基于退化的训练增强和评估时损坏。在Deepfake-Eval-2024的过滤子集上,在FakeAVCeleb上训练的模型分别达到了0.504、0.538和0.458的基线AUROC值。JPEG50训练增强将这些值分别提高到0.691、0.605和0.570,而所有三种准确率均下降。这些结果说明了为什么训练干预、评估损坏和依赖指标的结果应在共同流水线中保持区分。贡献在于集成了音视频处理、可互换检测器和可配置的实验工作流,并辅以实证案例研究。研究结果强调了跨数据集检测的挑战以及排名和分类指标提供的互补信息。
英文摘要
Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.