发表机构
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics; College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics; School of Computer Science and Technology, Harbin Institute of Technology(南京航空航天大学计算机科学与技术学院; 南京航空航天大学人工智能学院; 哈尔滨工业大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建了含三种生成协议的MLLM生成图像检测基准数据集,评估现有检测器性能并提出SAP-DSP基线框架,验证了其检测稳定性。
AI 中文摘要
近年来,GPT Image2、Nano Banana2等多模态大语言模型(MLLM)生成图像的逼真度快速提升。与早期生成模型相比,当前模型在文本渲染方面取得了明显进步,能够生成与真实世界应用场景高度相似的高质量图像。当前MLLM增强的生成能力对AI生成图像检测构成了日益严峻的挑战,检测工作不再局限于识别早期生成器留下的明显人工痕迹,而是需要针对新一代生成内容构建系统且逼真的基准数据集。然而,大多数现有基准仍围绕早期生成模型构建,无法充分评估高质量、多形式生成图像带来的取证挑战。为填补这一空白,本文构建了用于检测MLLM生成图像的基准数据集,该基准涵盖多个真实应用场景,并采用三种生成协议模拟直接生成、基于参考的重建及局部编辑。基于此基准,我们评估了检测器从传统场景到MLLM生成图像的性能退化情况,分析了三类样本的假阳性率与假阴性率,揭示了现有方法的失效模式。我们进一步提出结构人工痕迹先验引导的双流提示框架(SAP-DSP)作为强基线,该框架采用双流提示学习与结构感知路由融合以改进表示学习。大量实验表明,所提出的基准暴露了现有检测器在高质量生成图像上的性能退化,而SAP-DSP在该基准上实现了更稳定的检测结果。我们的代码与数据集公开于此https URL。
英文摘要
The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years. Compared with early generative models, current models have made clear progress in text rendering. They can produce high-quality images that closely resemble real-world application scenarios. The enhanced generation capabilities of current MLLMs pose increasingly severe challenges to AI-generated image detection. Detection is no longer limited to identifying obvious artifacts left by early generators. Instead, it requires systematic and realistic benchmarks for the new generation of generated content. However, most existing benchmarks are still built around early generative models and cannot fully evaluate the forensic challenges introduced by high-quality and multi-form generated images. To address this gap, this paper constructs a benchmark dataset for detecting images generated by MLLMs. The benchmark covers several realistic application scenarios and adopts three generation protocols to simulate direct generation, reference-based reconstruction, and local editing. Based on this benchmark, we evaluate detector degradation from traditional scenarios to MLLM-generated images and analyze false positive rates and false negative rates across three sample types, revealing the failure modes of existing methods. We further propose a structural-artifact-prior-guided dual-stream prompt framework (SAP-DSP) as a strong baseline. SAP-DSP uses dual-stream prompt learning and structure-aware routing fusion to improve representation learning. Extensive experiments show that the proposed benchmark exposes the performance degradation of existing detectors on high-quality generated images, while SAP-DSP achieves more stable detection results on this benchmark. Our code and dataset are publicly available at https://github.com/xbrainnet/SAP-DSP.