GAP-SAM:用于通用AI生成图像篡改定位的全局伪影先验
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization
浏览论文内容
中文总结 AI 辅助
本研究针对AI生成图像篡改定位的分布外性能问题,提出GAP-SAM方法,通过构建COCO-ControlNet对齐语义与几何,利用全局伪影标记抑制语义边界捷径,在六个数据集上平均像素F1达79.8,优于现有方法且在各类图像退化下表现最佳。
中文摘要 AI 辅助
AI生成图像篡改定位用于识别被编辑的像素,但其分布外(OOD)性能滞后于图像级检测,部分原因是像素级监督将取证证据与特定数据集的掩码几何结构及语义边界纠缠在一起。为将图像级分布对齐扩展至定位任务,我们构建了COCO-ControlNet,使用源图像的Canny边缘和深度图来对齐语义与几何,提升了多个定位器的OOD性能。然而更紧密的Mask-VAE重建对齐(Mask-VAE)的表现不及COCO-ControlNet,表明VAE重建伪影难以迁移至局部扩散修复伪影。我们还发现了「边界粘附」现象:微调后的分割模型会将预测结果贴合至语义对象轮廓,而非真实的篡改边界。这些发现催生了GAP-SAM,它将图像及其冻结的VAE重建编码为全局伪影标记,并在像素解码前通过零门控FiLM将其注入SAM3的特征金字塔。该标记不指定空间区域,而是调节密集解码以在保留定位能力的同时抑制语义边界捷径。在六个数据集上,GAP-SAM的平均像素F1值为79.8,比最强的现有方法高出12.6个百分点;在JPEG压缩、高斯模糊和调整大小的所有测试强度下,其表现均为最优。
英文摘要
AI-generated image manipulation localization identifies edited pixels, but its OOD performance lags behind image-level detection partly because pixel supervision entangles forensic evidence with dataset-specific mask geometry and semantic boundaries. Extending image-level distribution alignment to localization, we construct COCO-ControlNet with source-image Canny edges and depth maps to align semantics and geometry, improving OOD performance across multiple localizers. Yet tighter Mask-VAE Reconstruction Alignment (Mask-VAE) underperforms COCO-ControlNet, showing that VAE reconstruction artifacts transfer poorly to local diffusion-inpainting artifacts. We also identify \emph{boundary adhesion}, where fine-tuned segmentation models snap predictions to semantic object contours rather than true manipulation boundaries. These findings motivate GAP-SAM, which encodes an image and its frozen VAE reconstruction into a global artifact token and injects it into SAM3's feature pyramid via zero-gated FiLM before pixel decoding. Without prescribing a spatial region, this token modulates dense decoding to preserve localization while suppressing semantic-boundary shortcuts. Across six datasets, GAP-SAM averages 79.8 Pixel-F1, outperforming the strongest prior method by 12.6 points. It also performs best at every tested severity of JPEG compression, Gaussian blur, and resizing.