arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PatchHead:学习空间块证据以实现可泛化的AI生成图像检测

PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection

Shengbo Qi, Hongyi Fang, Benjia Zhou, Rui Mao

arXiv 2608.09223首次发表:更新:

AI 中文总结

针对AI生成图像检测器泛化能力差的问题,提出保留DINO块token二维结构的PatchHead,仅优化少量参数,在9个跨数据集基准上表现优异,提升了检测准确率并降低了域差异。

AI 中文摘要

AI生成图像检测器在训练和测试图像来自不同生成器或数据集时泛化能力较差。尽管像DINO这样的视觉基础模型能生成丰富的空间表示,但现有检测器通常仅使用全局聚合的CLS token对图像进行分类。我们假设,将DINO特征全局聚合为单个CLS token会掩盖空间分布的生成痕迹。为验证这一假设,我们提出PatchHead,这是一种轻量型空间聚合头,它保留DINO块token的二维结构,并整合相邻区域的证据。训练期间,我们冻结预训练的DINO主干,仅优化插入的LoRA适配器、PatchHead和辅助投影头。在涵盖人工整理和野外场景的9个跨数据集基准测试中,PatchHead在7个数据集上排名第一,在剩余2个数据集上排名第二。它将此前最强方法的平均平衡准确率从91.6%提升至94.6%(提升3.0个百分点),将最坏情况准确率从82.4%提升至89.4%(提升6.9个百分点),同时仅增加8.6%的可训练参数和0.08%的额外FLOPs。进一步的定性分析表明,PatchHead(i)减少了类条件域差异,(ii)将表示从内容主导的显著性转向空间分布的真实性证据。这些观察共同为以下现象提供了表示层面的解释:相较于基于单个CLS的全局表示,空间块聚合在不同生成器和数据集间的迁移更可靠。我们的代码和模型将在论文接收后公开。

英文摘要

AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spatial representations produced by vision foundation models like DINO, existing detectors typically classify images using only the globally aggregated CLS token. We hypothesize that globally aggregating DINO features into a single CLS token obscures spatially distributed generation traces. To test this hypothesis, we introduce PatchHead, a lightweight spatial aggregation head that preserves the two-dimensional organization of DINO patch tokens and integrates evidence across neighboring regions. During training, we freeze the pretrained DINO backbone and optimize only the inserted LoRA adapters, PatchHead, and auxiliary projection head. Across nine cross-dataset benchmarks spanning manually curated and in-the-wild settings, PatchHead ranks first on seven datasets and second on the remaining two. It improves the strongest prior method from 91.6% to 94.6% in average balanced accuracy (+3.0 points) and raises the worst-case accuracy from 82.4% to 89.4% (+6.9 points), while introducing only 8.6% more trainable parameters and 0.08% additional FLOPs. Further qualitative analysis suggests that PatchHead (i) reduces class-conditional domain discrepancy, and (ii) redirects the representation from content-dominated saliency toward spatially distributed authenticity evidence. Together, these observations provide a representation-level account of why spatial patch aggregation transfers more reliably across generators and datasets than a single CLS-based global representation. Our code and models will be made available upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑