arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37937cs.CV

看得更细:面向AI生成图像检测的块级监督

Look Closer: Patch-wise Supervision for AI-Generated Image Detection

  • Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)
  • ShanghaiTech University(上海科技大学)
  • Anhui University(安徽大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao

AI总结:

本研究提出块级监督方法,通过共享骨干网络对图像块分类并在推理时平均概率,无需额外模块,在GenImage上四种骨干网络均提升准确率,为AI生成图像检测提供了简单有效的方案。

AI中文摘要:

一个检测器需要看到图像的多少内容?即使小块RGB区域几乎不揭示整个场景的信息,它们仍可能保留图像合成过程的有用证据。受单块检测的启发,我们研究了块级监督:一个共享骨干网络对显式裁剪的图像块进行分类,每个裁剪块获得自己的损失,仅在推理时对块概率进行平均。该方法既不需要手工设计的残差滤波,也不需要学习的图像级融合模块。实验涵盖了单块选择、多个生成器集合以及四种CNN和Transformer骨干网络。在GenImage上,所报告的块级变体在所有四种骨干网络上相比其整图对应版本均提高了平均准确率。对监督粒度、源分辨率、裁剪尺寸和推理覆盖范围的比较进一步刻画了该方法的特点,而后处理测试和困难图像评估则揭示了其局限性。历史实验包括基于评估的模型选择,因此其分数不作为统一选择的排行榜比较呈现。总体而言,本研究将显式局部输入和块级监督确定为一种简单而有效的组合,用于研究可泛化的AI生成图像检测。

英文摘要:

How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.

补充信息

↑