发表机构
Rensselaer Polytechnic Institute; IBM Research(伦斯勒理工学院; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究能否用强化学习而非显式推理监督训练视觉-语言模型推理AI图像编辑,提出基于GRPO的框架,引入eff-IoU指标,在多数据集上实现与SOTA相当的检测定位性能。
AI 中文摘要
检测和定位AI篡改的图像对于值得信赖的AI至关重要,但现代生成模型使得这类篡改越来越难以识别。传统的二分类器可以检测图像篡改,但缺乏可解释性和泛化能力。视觉-语言模型(Vision-Language Models, VLMs)凭借其强大的视觉理解和推理能力成为有前景的替代方案;然而,现有方法通常依赖带有精心设计解释的监督微调,而非利用其内在的推理能力。本研究探究是否可以通过强化学习(Reinforcement Learning, RL)而非显式的推理监督来训练VLMs,使其能够对AI生成的图像编辑进行推理。受Group Relative Policy Optimization(GRPO)成功的启发——GRPO是一种通过在给出最终答案前生成思考轨迹来激励模型推理的RL技术,我们提出一种基于GRPO的训练框架,该框架利用简单的准确率和格式奖励。给定输入图像后,模型会生成结构化推理轨迹并预测图像是否被篡改;随后,轻量级分割模型在推理输出的引导下生成像素级定位掩码。在多个图像篡改数据集上开展的实验表明,尽管所需监督强度显著更低,我们的方法仍达到了与最先进图像伪造检测器相当的检测和定位性能。我们引入effective intersection over union(eff-IoU)这一统一指标,用于联合评估检测与定位效果。这些结果表明,强化学习为训练VLMs对AI生成内容进行推理提供了一种有效且可扩展的机制。
英文摘要
Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Language Models (VLMs) offer a promising alternative due to their strong visual understanding and reasoning capabilities; however, existing approaches typically rely on supervised finetuning with curated explanations rather than exploiting their inherent reasoning capabilities. In this work, we investigate whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision. Motivated by the success in Group Relative Policy Optimization (GRPO), an RL technique that incentivizes the model to reason by generating thinking traces prior to giving the final answer, we propose a GRPO-based training framework that utilizes simple accuracy and format rewards. Given an input image, the model produces a structured reasoning trace and predicts whether the image has been tampered with. A lightweight segmentation model is then guided by the reasoning output to generate pixel-level localization masks. Experiments across multiple image manipulation datasets demonstrate that our approach achieves competitive detection and localization performance compared to state-of-the-art image forgery detectors, despite requiring substantially weaker supervision. We introduce effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization. These results suggest that reinforcement learning provides an effective and scalable mechanism for teaching VLMs to reason about AI-generated content.