通过噪声残差聚类实现有效的合成图像检测
Effective Synthetic Image Detection via Noise Residual Clustering
AI总结:
针对合成图像检测问题,提出一种无训练方法,先提取噪声残差指纹,再经视觉Transformer提取多尺度特征并融合,用少量真实图像样本初始化聚类中心,在四个基准数据集上平均准确率达82.2%,泛化能力优于现有检测器。
AI中文摘要:
生成式人工智能的快速发展使合成图像极为逼真,带来诸如错误信息和欺诈等安全威胁。以被动和盲图像认证方式检测合成图像具有重要意义。大多数现有检测器依赖大量标记数据集的监督训练,成本高且对未知生成模型性能下降。为此提出一种无训练检测方法,先由预训练的Noiseprint++模型提取噪声残差指纹,再用冻结的视觉Transformer提取多尺度特征并自适应加权融合,仅用少量真实图像样本初始化无监督K-Means聚类中心来区分真实与合成图像。在四个基准数据集上的广泛评估表明,该方案平均准确率达82.2%,在泛化能力上优于现有检测器,在流行的扩散型合成图像上性能优越,且通过消融研究验证了各模块的有效性。
英文摘要:
The rapid advancement of generative artificial intelligence (AI) has made synthetic images remarkably realistic, posing security threats such as misinformation and fraud. It is significant to detect the synthetic image in the manner of passive and blind image authentication. Most existing detectors rely on supervised training with large labeled datasets, leading to high costs and degraded performance on unknown generative models. To attenuate such deficiencies, we propose a training-free detection method. Specifically, noise residual fingerprints are first extracted by a simple yet effective pre-trained Noiseprint++ model. Then multi-scale features are further extracted from such residual by a frozen Vision Transformer (ViT), followed by adaptive weighted fusion. Only a few real image samples are used needed to initialize the clustering centers for unsupervised K-Means, distinguishing real and synthetic images without training. Extensive evaluations on four benchmark datasets show that our proposed scheme achieves an average accuracy of 82.2%, outperforming the state-of-the-art detectors on generalization ability. Superior performance is gained on the popular diffusion type of synthetic images, and the effectiveness of each module is validated by ablation studies. Source code will be publicly available at https://github.com/multimediaFor/NoiseCluSID.