法医双胞胎:面向AI生成图像取证的自我监督残差学习
Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics
- Universidad Autónoma de Madrid(马德里自治大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出法医双胞胎框架,利用自我监督残差学习在真实图像上训练,通过抑制宏观内容提取采集指纹,实现零样本AI生成图像检测与源归因,分别达到97.99% AUC和56.61%准确率。
AI中文摘要:
AI生成图像的检测器通常使用它们必须捕获的所有生成式AI架构的样本进行训练,一旦新架构出现就会难以应对。近期方法探索了自我监督预训练作为替代方案,但标准框架与取证任务相悖,例如其数据增强会覆盖图像形成的微观统计特征。本文提出了法医双胞胎(Forensic Twins),一种自我监督残差学习(SSRL)框架,其前置任务抑制宏观内容可用性。每张图像通过冻结的、现成的取证残差提取器映射,从中提取两个空间上不相交的裁剪块。两个视图不共享任何像素,保留最少的语义结构用于对齐,使得冗余减少目标以占主导地位的共同信号为主:图像采集管道的平稳指纹。此外,法医双胞胎仅使用真实图像训练;在任何阶段都不观察AI生成的图像。实验表明,法医双胞胎以56.61%的准确率归因AI生成器来源,比之前的零样本最先进方法高出6.13%,且延迟降低375倍。我们还证明,仅使用从法医双胞胎提取的真实图像嵌入离线拟合高斯混合模型(GMM),可将其转变为最先进的零样本检测器,在27个未见过的AI生成器(包括GAN、扩散模型和商业系统)上达到97.99%的AUC。代码、权重和精确划分将公开提供。
英文摘要:
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretext task suppresses macroscopic content availability. Each image is mapped through a frozen, off-the-shelf forensic residual extractor, from which two spatially disjoint crops are drawn. Sharing no pixel, the two views retain minimal semantic structure to align, leaving a redundancy-reduction objective with a predominant common signal: the stationary fingerprint of the image acquisition pipeline. Additionally, Forensic Twins is trained exclusively on real images; no AI-generated image is observed at any stage. Experiments show that Forensic Twins attributes AI generator sources with 56.61% accuracy, i.e., 6.13% above the previous state-of-the-art zero-shot method at 375x lower latency. We also demonstrate that fitting a Gaussian Mixture Model (GMM) offline using only the real image embeddings extracted from Forensic Twins turns it into a state-of-the-art zero-shot detector, reaching 97.99% AUC across 27 unseen AI generators, including GANs, diffusion models and commercial systems. Code, weights and exact splits will be made publicly available