ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
ERF-BA-TFD+: 一种用于音频视觉深度伪造检测的多模态模型
专题命中 音视频/视觉语言融合 :audio-visual fusion(abstract)
AI总结 ERF-BA-TFD+通过结合增强接收场和音频视觉融合,提出了一种多模态深度伪造检测模型,在DDL-AV数据集上实现了最先进的检测性能。
Comments The paper is withdrawn after discovering a flaw in the theoretical derivation presented in Section Method. The incorrect step leads to conclusions that are not supported by the corrected derivation. We plan to reconstruct the argument and will release an updated version once the issue is fully resolved