arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Swin与EfficientNet结合:基于GAN的人脸取证轻量级架构

Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics

Sejuti Basu, Ashima Sood, Vijay Kumar, Sahil Sharma

arXiv 2609.01749首次发表:更新:

发表机构

International Institute of Information Technology, Bangalore; Ulster University; Dr B R Ambedkar National Institute of Technology Jalandhar(国际信息技术学院(班加罗尔); 阿尔斯特大学; 阿姆倍伽尔博士国家技术学院贾朗达尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出结合EfficientNet-B0卷积处理与Swin Transformer后端的混合架构,在14万张人脸数据集上实现99%准确率,高效检测GAN生成的合成人脸,性能优于纯Swin变体及仅CNN的基线。

AI 中文摘要

现代生成模型如GAN、扩散架构和自回归系统,现已能生成与真实照片几乎无法区分的人脸图像。这种能力使得伪造图像的检测愈发困难,引发了关于身份盗窃、欺诈和虚假信息传播的严重担忧。本研究聚焦于支撑许多以人脸为中心的深度伪造的GAN生成合成人脸,仅利用图像分析研究高效检测方法。现有检测系统严重依赖卷积神经网络(CNN)或全局视觉Transformer:CNN擅长识别基于纹理的局部特征,但在更广泛的上下文理解上存在困难;传统视觉Transformer(ViT)模型可有效捕捉长程结构,但需要大量计算资源。本研究探索基于Swin-Transformer的架构,包含三种实现方式:从头训练的紧凑Swin Transformer、针对二分类任务适配的ImageNet-1K预训练Swin-Tiny和Swin-Small模型,以及结合EfficientNet-B0的卷积处理与Swin Transformer后端的新型混合模型。我们使用包含14万张真实与伪造人脸的数据集评估所有模型,该数据集包含StyleGAN生成的伪造人脸,以及来自Flickr和DFDC的真实图像,训练、验证和测试集划分均衡。在5000张测试图像上,EfficientNetB0+Swin混合模型取得了99%的准确率和99.44%的召回率,在该数据集上的表现优于纯Swin变体和此前仅基于CNN的基线模型。我们的结果表明,将分层CNN特征与移位窗口自注意力结合,为检测GAN生成的合成人脸提供了一种高效且计算轻量的方法。

英文摘要

Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult, raising serious concerns about identity theft, fraud, and misinformation campaigns. Our research focuses specifically on GAN-generated synthetic faces, which underpin many face-centric deepfakes, and investigates efficient detection approaches using image analysis alone. Existing detection systems rely heavily on either convolutional neural networks (CNNs) or global vision transformers. While CNNs excel at identifying texture-based local features, they struggle with broader contextual understanding. Traditional Vision Transformer (ViT) models can capture long-range structures effectively, but demand substantial computational resources. Our work explores Swin-Transformer-based architectures across three implementations: a compact Swin Transformer trained from the ground up, ImageNet-1K pre-trained Swin-Tiny and Swin-Small models adapted for binary classification, and a novel hybrid combining EfficientNet-B0's convolutional processing with a Swin Transformer backend. We evaluated all models using the 140K Real and Fake Faces dataset, which includes StyleGAN-generated fake faces alongside authentic images from Flickr and DFDC, with balanced splits for training, validation, and testing. The EfficientNetB0+Swin hybrid achieved 99% accuracy and a 99.44% recall on 5,000 test images, outperforming both pure Swin variants and a previous CNN-only baseline on this dataset. Our results suggest that combining hierarchical CNN features with shifted-window self-attention provides an efficient and computationally lightweight method for detecting GAN-generated synthetic faces.

Comments12 pages, 2 figures, 1 table. Presented at the International Conference on Computational Techniques in Data Science (IICTDS 2025), online. Proceedings forthcoming

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑