发表机构
Mohamed bin Zayed University of Artificial Intelligence; University College London; Weizmann Institute of Science(穆罕默德·本·扎耶德人工智能大学; 伦敦大学学院; 魏茨曼科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究现代视觉语言模型中像素级图像篡改检测的域泛化,提出基于平衡小批量采样和后期注入策略的简单训练框架,大幅提升平均gIoU和cIoU,增强了篡改定位和分布外鲁棒性。
AI 中文摘要
现代视觉语言模型(VLMs)显著提升了图像生成和编辑能力,使得像素级图像篡改检测在跨模型和分布外转移下愈发重要且具有挑战性。本文研究了像ChatGPT、Gemini、Qwen-Image等现代VLMs中像素级图像篡改检测的域泛化,旨在学习在不同VLM生成的操纵分布中保持稳健的篡改定位模型。我们提出了一个基于两种实用策略的简单而有效的域泛化训练框架。首先,引入平衡小批量采样方案,在每个小批量中策略性地采样篡改和真实图像,防止对操纵伪像或干净图像先验的偏差优化,避免训练崩溃,确保每个优化步骤接收适当采样的梯度信号。其次,采用简单的后期注入策略,探测器先在大规模基础数据上训练至稳定收敛,然后接触来自新兴VLM分布的少量新选择的支持数据,提高适应性而不过度拟合有限的新域。这些组件共同提供了一个简单而强大的方法来改进现代VLMs中像素级篡改定位和分布外鲁棒性。尽管概念简单,我们的框架在GPT-Images-2.0、Gemini-3.1、FLUX.2和Seedream 4.5的分布外VLMs上,平均gIoU和cIoU分别比先前的最优方法PIXAR有26.1%和26.8%的大幅相对提升。我们的代码可在该https URL获取。
英文摘要
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two practical strategies. First, we introduce a balanced minibatch sampling scheme that strategically samples tampered and real images in each minibatch, preventing biased optimization toward either manipulated artifacts or clean-image priors and avoiding training collapse, ensuring that each optimization step receives proper sampled gradient signals. Second, we adopt a simple late-injection strategy, where the detector is first trained on large-scale base data until stable convergence, and then exposed to a small amount of newly selected supporting data from emerging VLM distributions, improving adaptability without overfitting to limited new domains. Together, these components provide a simple yet strong recipe for improving pixel-level tampering localization and OOD robustness across modern VLMs. Despite the conceptual simplicity, our framework outperforms the prior state-of-the-art PIXAR by a large margin of 26.1% and 26.8% relative improvement in average gIoU and cIoU, respectively, across OOD VLMs of GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5. Our code is available at https://github.com/VILA-Lab/PIXAR-DG
CommentsOur code is available at https://github.com/VILA-Lab/PIXAR-DG