arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25842cs.CVcs.CR

对抗性深度伪造生成与基于净化的对抗性检测研究

Adversarial Deepfake Generation and an Investigation of Purification-Based Adversarial Detection

Junghyun Kim, Seunghyun Kim, Jiyoung Woo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究参与ImageCLEF 2026深度伪造检测与生成任务,图像生成采用FLUX.1-dev等方法,检测用SigLIP+DINOv2等组合。还自主研究基于净化的对抗性检测,发现特定信号可有效区分对抗与干净输入,反驳相关假设并揭示质量悬崖。

中文摘要 AI 辅助

本文描述了“Go To Germany”团队参与ImageCLEF 2026深度伪造检测与生成任务的情况。对于图像生成任务,采用FLUX.1-dev和PuLID进行身份保留人脸合成,并结合多模型PGD对抗攻击同时针对12个检测器。该方法在躲避组织者检测器方面达到90%,参与者检测器方面为57.6%,最终生成分数为0.4170。对于图像检测任务,将SigLIP+DINOv2和GenD-DINOv3两个互补检测器组合成最大概率集成,在基线深度伪造上准确率达99.4%,但在真实图像上误报率高,最终检测分数为0.6986。此外,还对基于净化的对抗性检测进行了自主研究,比较了六个共享CLIP ViT-L/14骨干的检测器的三种检测信号家族。发现通过EFFORT检测器应用的中位数-3净化下的原始$|\Delta \text{logit}|$,在四种对抗源类型中以0.81-0.98的AUROC将对抗性输入与干净输入分开,这一发现反驳了简单的骨干保留假设,并揭示了在Q70处信号崩溃的尖锐JPEG质量悬崖。

英文摘要

This paper describes the participation of team "Go To Germany" in the ImageCLEF 2026 Deepfake Detection and Generation Task. For the image generation task, we employ FLUX.1-dev with PuLID for identity-preserving face synthesis, combined with a multi-model PGD adversarial attack targeting 12 detectors simultaneously (DiffJPEG-in-loop, MI/DI/EoT, adaptive weighting, two-stage warm-start). Our approach achieved 90% evasion against organizer detectors and 57.6% against participant detectors, with a final generation score of 0.4170. For the image detection task, we combine two complementary detectors - SigLIP+DINOv2 for AI-generated images and GenD-DINOv3 for face manipulations - in a max-probability ensemble, achieving 99.4% accuracy on baseline deepfakes but suffering from high false-positive rates on real images, resulting in a final detection score of 0.6986. Beyond the official submission, we conducted a self-initiated investigation of purification-based adversarial detection, comparing three families of detection signals across six detectors that share a CLIP ViT-L/14 backbone. We find that raw $|Δ\text{logit}|$ under median-3 purification, applied through the EFFORT detector, separates adversarial inputs from clean inputs with AUROC 0.81-0.98 across four adversarial source types - a finding that refutes the simple backbone-preservation hypothesis and exposes a sharp JPEG-quality cliff at Q70 where the signal collapses.

发表机构

  • Soonchunhyang University(顺天乡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑