arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15196cs.CV

基于DINOv3的可泛化AI生成图像检测的锚点正则化适配方法

Anchor-Regularized Adaptation for Generalizable AI-Generated Image Detection with DINOv3

  • Sungkyunkwan University(成均馆大学)
  • University of Naples Federico II(那不勒斯费德里科二世大学)
  • Secure Machines Lab Inc(Secure Machines Lab公司)

机构由 AI 辅助整理,请以论文原文为准。

Hyeongjun Choi, Juhun Lee, Davide Cozzolino, Luisa Verdoliva, Simon S. Woo

AI总结:

针对AI生成图像检测中混合对齐与未对齐数据易扭曲预训练表示的问题,提出锚点正则化适配(ARA)方法,结合低秩适配与冻结锚点分类器,在9个基准测试中取得最优性能。

AI中文摘要:

近期AI生成图像检测领域的研究表明,精心设计的训练数据对齐可通过消除虚假关联提升模型泛化能力。然而,在冻结的DINOv3表示上进行线性探测,即便在未对齐数据集上训练,也能取得极强的性能。受此结果启发,我们分析了这种泛化的潜在原理与局限性。研究发现,冻结的DINOv3表现良好,是因为其决策依赖于能忠实表征真实图像空间的特征;同时,其最终层在捕捉可通过对齐训练数据强化的细微像素伪影线索方面效果较差。我们进一步观察到,在适配过程中直接混合对齐与未对齐数据,虽能提升对这类线索的敏感性,但会以扭曲预训练表示为代价,进而限制泛化能力。为解决该问题,我们提出锚点正则化适配(Anchor-Regularized Adaptation, ARA)方法:应用低秩适配(Low-Rank Adaptation)以捕捉像素级伪影,同时利用冻结的锚点分类器避免偏离原始表示结构。这使模型能利用像素伪影线索,且不牺牲泛化能力。我们的方法在9个多样化且具有挑战性的基准测试中取得了最先进的性能,表明ARA能从未对齐与对齐数据中获取互补监督,实现更有效的检测。

英文摘要:

Recent works in AI-generated image detection have shown that careful training data alignment can improve generalization by removing spurious correlations. However, linear probes on frozen DINOv3 representations achieve remarkably strong performance even when trained on misaligned datasets. Motivated by this result, we analyze the underlying rationale and the limits of this generalization. We find that frozen DINOv3 performs well because its decisions rely on features that faithfully represent the space of authentic images. At the same time, its final layer is less effective at capturing the subtle pixel-artifact cues that can be emphasized by aligned training data. We further observe that naively mixing aligned and misaligned data during adaptation improves sensitivity to such cues but at the cost of distorting the pre-trained representation, limiting generalization. To address this issue, we propose Anchor-Regularized Adaptation (ARA). We apply Low-Rank Adaptation to capture pixel-level artifacts while leveraging a frozen anchor classifier to avoid deviations from the original representation structure. This allows the model to exploit pixel-artifact cues without sacrificing generalization. Our method achieves state-of-the-art performance on nine diverse and challenging benchmarks, indicating that ARA enables complementary supervision from misaligned and aligned data for more effective detection.

↑