发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成图像检测泛化性不足的问题,提出RED框架,利用多尺度重建演化中的token可预测性反转捕获取证线索,在六个基准上达到92.5%平均准确率。
AI 中文摘要
图像生成器的快速演进要求取证线索能够泛化到已知生成机制之外。现有检测器通常依赖静态图像表示或端点重建差异,忽视了中间重建阶段的演化过程。我们观察到,真实图像与生成图像的相对token可预测性在不同重建尺度上可能发生反转,这表明中间阶段可能暴露了端点比较所忽略的取证证据。基于这一观察,我们提出RED(重建演化动力学)框架,从粗到细的重建演化中捕获可迁移的取证线索。据我们所知,RED是首个利用尺度级token可预测性来引导跨中间重建状态取证证据聚合的框架。它将冻结的多尺度VQ-VAE产生的重建轨迹表示在冻结的CLIP编码器的共享特征空间中。为了将观察到的可预测性变化与视觉证据联系起来,RED从冻结的VAR模型提供的尺度级token负对数似然中学习图像自适应阶段权重。随后,跨阶段证据聚合模块联合建模原始图像表示和加权重建特征,通过沿重建轨迹的交互捕获互补的取证线索。在六个多样化基准上的实验表明,RED在评估方法中实现了最高的平均准确率92.5%和平均精确率97.5%。进一步评估显示其对常见图像退化具有强鲁棒性,支持重建演化用于泛化AI生成图像检测的价值。代码将在论文被接收后公开提供。
英文摘要
The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5\% and average precision of 97.5\% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.