arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向以人为中心场景中AI生成图像检测的视觉证据定位与解释

Grounding and Explaining Visual Evidence for AI-Generated Image Detection in Human-Centric Scenes

Kun Guo, Yuzhou Yang, Haoyue Wang, Qichao Ying, Sheng Li, Zhenxing Qian

arXiv 2608.01988首次发表:更新:

AI 中文总结

针对以人为中心场景AI生成图像检测的证据定位与解释局限,提出HAVE数据集及PAVE框架,实现真伪预测、证据定位与解释生成,实验表现优异。

AI 中文摘要

图像生成模型的快速发展催生了可解释的AI生成图像检测方法,这类方法不仅需要判断图像真伪,还需提供支持性视觉证据。现有方法可能存在生成的解释与局部证据区域不一致的问题,削弱了真伪判定解释的可靠性;同时,现有基准对生成图像中普遍存在的多样化以人为中心场景覆盖有限。为解决这些局限,我们研究以人为中心场景中带有定位和可解释视觉证据的真伪检测问题,提出HAVE(Human-centric AI-generated Visual Evidence)数据集,该多样化以人为中心数据集包含来自10种近期生成器的4万张真实图像和3.9万张AI生成图像,涵盖8类证据的10.6万条局部证据实例,每条实例均标注边界框和区域对齐解释。我们进一步提出PAVE(Perception-Aware Visual Evidence)框架,该框架联合执行真伪预测、视觉证据定位和区域对齐解释生成,采用裁判引导的对齐奖励评估区域-解释一致性和证据有效性,同时结合感知感知正则化,对比原始图像与随机掩码图像的令牌级预测,以促进模型依赖视觉输入。在HAVE数据集及外部数据集上的实验表明,该方法在真伪检测、视觉证据定位和解释质量方面表现出色,代码和数据将在发表后发布。

英文摘要

Rapid advances in image generation models call for interpretable AI-generated image detection methods that not only determine authenticity but also provide supporting visual evidence. Existing approaches may produce inconsistencies between generated explanations and localized evidence regions, undermining the reliability of explanations for authenticity decisions. Meanwhile, existing benchmarks provide limited coverage of the diverse human-centric scenes prevalent in generated imagery. To address these limitations, we investigate authenticity detection with grounded and explainable visual evidence in human-centric scenes. We present HAVE (Human-centric AI-generated Visual Evidence), a diverse human-centric dataset comprising 40K real and 39K AI-generated images from 10 recent generators, with 106K localized evidence instances across 8 evidence categories, each annotated with a bounding box and a region-aligned explanation. We further propose PAVE, a Perception-Aware Visual Evidence framework that jointly performs authenticity prediction, visual evidence grounding, and region-aligned explanation generation. PAVE employs a judge-guided alignment reward to assess region--explanation consistency and evidence validity, together with perception-aware regularization that contrasts token-level predictions between original and randomly masked images to promote reliance on visual input. Experiments on HAVE and external datasets demonstrate strong performance in authenticity detection, visual evidence grounding, and explanation quality. Code and data will be released upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑