AI 中文总结
针对深度伪造检测问题,提出基于快速傅里叶变换的MSCA-FFT框架,结合Xception空间分支与FFT频率分支,经变压器编码器细化、交叉注意力融合后由MLP分类,性能优于DCT方法及基线模型,且频率分支能提供互补线索并给出可解释证据。
AI 中文摘要
深度伪造的产生引发了人们对数字媒体真实性、错误信息、身份欺诈和公众信任等方面的日益关注。近期研究表明,结合空间和频率特征比单独使用能带来更强的检测效果。本文提出了MSCA-FFT,一种基于快速傅里叶变换(FFT)的多尺度交叉注意力框架用于图像级深度伪造检测。该模型将部分微调的Xception空间分支与基于FFT的频率分支相结合。频率分支通过浅卷积层处理对数缩放的FFT幅度谱,避免了基于离散余弦变换(DCT)的流程中使用的逆频率到图像重建。空间和频率表示由变压器编码器细化,通过交叉注意力融合,并传递给MLP分类器进行真假预测。实验结果表明,MSCA-FFT的性能始终高于基于DCT的先进空间频率融合方法和比较的基线模型。消融研究进一步表明,基于FFT的频率分支在与空间特征融合时提供了互补的光谱线索。此外,基于FFT的频率分析和Grad-CAM/LIME解释在包括眼睛、嘴巴、鼻子和面部边界等对操纵敏感的面部区域周围显示出一致的证据。
英文摘要
Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust. Recent studies show that combining spatial and frequency features leads to stronger detection results than using independently. This paper presents MSCA-FFT, a Fast Fourier Transform (FFT)-based multi-scale cross-attention framework for image-level deepfake detection. The model combines a partially fine-tuned Xception spatial branch with an FFT-based frequency branch. The frequency branch processes the log-scaled FFT magnitude spectrum through shallow convolutional layers, avoiding inverse frequency-to-image reconstruction used in DCT-based pipelines. The spatial and frequency representations are refined by transformer encoders, fused through cross-attention, and passed to an MLP classifier for real/fake prediction. Experimental results show that MSCA-FFT achieves consistently higher performance than the DCT-based state-of-the-art spatial-frequency fusion method and the compared baseline models. The ablation study further indicates that the FFT-based frequency branch provides complementary spectral cues when fused with spatial features. In addition, FFT-based frequency analysis and Grad-CAM/LIME explanations show consistent evidence around manipulation-sensitive facial regions, including the eyes, mouth, nose, and facial boundaries.