arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究基础模型为何适用于扩散生成图像检测

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

Davide Cozzolino, Giovanni Poggi, Luisa Verdoliva

arXiv 2608.12155首次发表:更新:

发表机构

University Federico II of Naples(那不勒斯费德里科二世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究基础模型用于扩散生成图像检测的原因,通过DDIM反演、频率交换及潜在空间分析,发现其利用中低频分布差异实现检测,为该类方法的可解释性提供了新方向。

AI 中文摘要

视觉基础模型近期已成为检测AI生成图像的强大特征提取器,在跨生成器的泛化性以及对常见图像降质的鲁棒性方面表现出色,但其有效性背后的原因却鲜为人知。本研究旨在探究基于基础模型的检测器利用何种线索区分真实图像与扩散生成图像,为此设计了一种基于DDIM反演的专用分析协议:给定一张真实图像,通过改变DDIM反演的深度生成一系列合成副本,尽管多数副本在语义上与真实参考图像一致,但由于扩散合成引入的细微痕迹,检测器的得分在这些副本间存在显著差异,表明其决策并非主要由语义失败驱动。通过频率交换分析进一步揭示,检测器利用的判别线索主要集中在中低频范围,而非生成模型相关伪影通常所在的高频范围;潜在空间分析则显示,再生图像的方差和有效维度降低,表明扩散模型未能完全复现真实数据的变异性。总体而言,研究结果表明,基于基础模型的检测器通过捕捉真实图像与扩散生成图像之间的非语义中低频分布差异取得成功,这些发现为这类检测器的鲁棒性与泛化性提供了新的见解,并为开发更具可解释性的取证方法指明了方向。

英文摘要

Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across generators and robustness to common image degradations. However, the reason behind their effectiveness is poorly understood. In this work, we investigate what cues are exploited by foundation-model-based detectors to distinguish real images from diffusion-generated ones. To this end, we design an ad hoc analysis protocol based on DDIM inversion. Given a real image we generate a sequence of synthetic copies by changing the depth of DDIM inversion. Even though most copies are semantically identical to the real reference, the detector score varies significantly across them due to subtle traces introduced by the diffusion synthesis, showing that its decision is not primarily driven by semantic failures. Through a frequency-swapping analysis, we further reveal that the discriminative cues exploited by the detectors are mainly localized in the low-to-mid frequency range, rather than only in the high-frequency range, as is the case for artifacts commonly associated with generative models. Finally, a latent-space analysis shows that regenerated images exhibit reduced variance and effective dimensionality, indicating that diffusion models do not fully reproduce the variability of real data. Overall, our results suggest that foundation-model-based detectors succeed by capturing non-semantic low-to-mid frequency distributional discrepancies between real and diffusion-generated images. These findings provide new insight into the robustness and generalization of such detectors and suggest directions for more interpretable forensic methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑