发表机构
Hefei University of Technology; University of Macau; University of Science and Technology of China(合肥工业大学; 澳门大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成图像检测中现有方法受上下文偏移影响的问题,提出相对块响应学习(PRL),通过三个目标优化,在多基准上显著提升检测性能。
AI 中文摘要
生成模型如今能够合成高度逼真的图像,同时增加了虚假信息和视觉伪造的风险。因此,检测AI生成图像变得愈发重要,可靠的检测器必须能泛化至未见过的生成器,并对现实场景中未见过的扰动保持鲁棒性。现有检测器通常在独立收集的真实图像和生成图像上训练,或在为缓解内容偏差而设计的对齐真实-生成图像对上训练。基于对齐对,近期方法通过用生成图像对应块替换真实图像的部分块来形成混合视图。然而,我们发现自注意力机制会让真实块和生成块产生交互,因此每个块的特征不再仅反映其自身来源。这种上下文偏移使得逐块来源标签成为不精确的目标。为此,我们提出相对块响应学习(Relative Patch Response Learning, PRL)。PRL不标记每个块,而是比较对齐对的两个混合视图中相同块,并从其块响应(即该块在两个视图间的分数变化)中学习:(i)为在上下文偏移下提供精确监督,相对响应目标将来源改变块的响应与来源未改变块的响应(后者仅响应偏移)进行对比;(ii)为偏移提供可靠参考,参考一致性目标使每组来源未改变块作为整体移动;(iii)由于两个视图包含的生成内容量不同,区域排名目标要求生成区域更大的视图具有更高的平均块分数。大量实验表明,PRL性能优异,在8个标准基准和3个现实场景基准的平均平衡准确率上,分别超过现有最佳方法4.3%和5.9%。
英文摘要
Generative models can now synthesize highly realistic images, simultaneously increasing the risks of misinformation and visual forgery. Therefore, detecting AI-generated images becomes more essential, and a reliable detector must generalize to unseen generators and stay robust to unseen perturbations in the wild. Existing detectors are typically trained on either independently collected real and generated images or aligned real-generated pairs designed to mitigate content bias. Building on aligned pairs, recent methods form a mixed view by replacing some patches of the real image with their generated counterparts. However, we find that self-attention lets real and generated patches interact, so the feature of each patch no longer reflects its own source alone. This contextual shift makes a per-patch source label an imprecise target. To this end, we propose Relative Patch Response Learning (PRL). Instead of labeling each patch, PRL compares the same patch across two mixed views of an aligned pair and learns from its patch response, the change of its score between the views. (i) To give precise supervision under the contextual shift, a relative response objective measures the responses of source-changed patches against those of source-unchanged patches, which respond to the shift alone. (ii) To provide a reliable reference for the shift, a reference coherence objective keeps each group of source-unchanged patches moving as a whole. (iii) Since the two views contain different amounts of generated content, an area ranking objective asks the view with the larger generated area to have a higher mean patch score. Extensive experiments demonstrate the superior performance of PRL, which surpasses the best prior methods by 4.3% and 5.9% in average balanced accuracy across eight standard and three in-the-wild benchmarks, respectively.
Comments19 pages, 6 figures, 13 tables