arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08639cs.CV

法证储备:激发图像伪造检测的潜在知识

Forensic Reserve: Eliciting Latent Knowledge for Image Forgery Detection

Jiahua Li, Zixu John, Tom Zhong, Fuping Wu, Tianhao Xu, Jianqing Zheng, Yuanhan Mo, Fei Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出储备引导激发(RGE)框架,利用预训练模型内部稀疏的法证组件作为储备,通过轻量级适配器激发潜在知识,仅用500张图像和0.2%参数即在多个伪造检测基准上取得竞争性能。

中文摘要 AI 辅助

随着生成图像变得越来越逼真,可靠的伪造检测对于维护视觉信息的可信度至关重要。然而,现有方法主要依赖任务特定的监督来调整视觉基础模型的表示,未能充分利用内部法证知识来指导检测。为解决这一局限,我们提出了储备引导激发(Reserve-Guided Elicitation, RGE)框架,该框架将预训练模型中稀疏的、对来源敏感的内部组件视为法证储备,并将其定位转化为轻量级适配的结构约束。具体而言,我们首先使用法证透镜(Forensic Lens, F-lens)将各层和令牌组的激活分解为独立组件,并根据它们在真实图像与生成图像之间的响应差异进行全局筛选,从而识别储备位点和方向。接下来,我们将选定的方向映射回隐藏状态空间以构建固定的储备子空间,并仅在识别出的位点插入法证储备适配器(Forensic Reserve Adapters, FRA)。最后,在固定骨干参数、先前拟合的参考分类器和子空间基的情况下,我们仅训练FRA系数映射,以生成受限于相应子空间的输入相关残差更新,从而强化现有的法证响应。仅使用500张标注训练图像和低于骨干0.2%的可训练参数预算,RGE在三个检测基准上取得了具有竞争力的性能,且无需针对目标基准进行适配。此外,RGE在涵盖自监督和视觉-语言预训练的八个编码器上持续优于相应的冻结检测器,激发了预训练视觉模型中广泛共享的潜在法证能力。

英文摘要

As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existing methods primarily rely on task-specific supervision to adapt vision foundation model representations, without fully exploiting internal forensic knowledge to guide detection. To address this limitation, we propose Reserve-Guided Elicitation (RGE), a framework that treats sparse, origin-sensitive internal components in pretrained models as a forensic reserve and translates their localization into structural constraints for lightweight adaptation. Specifically, we first use the Forensic Lens (F-lens) to decompose activations across layers and token groups into independent components and globally screen them by their response differences between real and generated images, identifying reserve sites and directions. Next, we map the selected directions back to hidden-state space to construct fixed reserve subspaces and insert Forensic Reserve Adapters (FRA) only at the identified sites. Finally, with the backbone parameters, previously fitted reference classifier, and subspace bases fixed, we train only the FRA coefficient maps to generate input-dependent residual updates constrained to the corresponding subspaces, strengthening existing forensic responses. Using only 500 labeled training images and a trainable parameter budget below 0.2% of the backbone, RGE achieves competitive performance across three detection benchmarks without target-benchmark adaptation. Furthermore, RGE consistently improves over the corresponding frozen detectors across eight encoders spanning self-supervised and vision-language pretraining, eliciting a latent forensic capacity broadly shared across pretrained vision models.

发表机构

  • University of Oxford(牛津大学)
  • Imperial College London(伦敦帝国理工学院)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑