arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成图像更易被遗忘:合成图像检测的机器遗忘视角

Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection

Jun Nie, Yonggang Zhang, Tongliang Liu, Yiu-ming Cheung, Bo Han, Xinmei Tian

arXiv 2608.00716首次发表:更新:

发表机构

University of Science and Technology of China; Hong Kong Baptist University; The Hong Kong University of Science and Technology; The University of Sydney(中国科学技术大学; 香港浸会大学; 香港科技大学; 悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对合成图像检测任务,发现大规模视觉模型在机器遗忘时对生成图像的特征退化更快,据此提出两种基于机器遗忘的检测方法,其性能优于传统方法,为合成图像检测建立了新范式。

AI 中文摘要

生成图像的鲁棒检测对于应对生成模型的滥用至关重要。现有方法主要依赖从人工标注的训练数据集中学习,限制了其对未见过分布的泛化能力。相比之下,在网络规模数据集上预训练的大规模视觉模型(LVMs)通过接触多样化分布展现出出色的泛化能力,为该任务提供了变革性范式。然而,我们的实验结果表明,以自然图像为主的数据预训练的LVMs能够有效捕捉自然图像和生成图像的特征,产生的损失相当低,因此在两者之间的判别能力有限。这引发了一个关键问题:LVMs在捕捉自然图像和生成图像特征时,何时以及如何表现出不同的行为?本研究揭示了一个见解:在机器遗忘过程中,LVMs表现出不同的遗忘动态,生成图像的特征退化速度比自然图像更快。受这种不同动态的启发,我们引入了两种检测方法:1)无数据检测,即修剪模型参数以在不访问数据的情况下诱导机器遗忘;2)数据驱动检测,即优化LVMs以遗忘与生成图像相关的知识。在各种基准上进行的大量实验表明,我们基于机器遗忘的方法优于传统检测方法。通过将检测任务重新定义为机器遗忘问题,本研究为生成图像检测建立了新范式。

英文摘要

Robust detection of generated images is critical to counter the misuse of generative models. Existing methods primarily depend on learning from human-annotated training datasets, limiting their generalization to unseen distributions. In contrast, large-scale vision models (LVMs) pre-trained on web-scale datasets exhibit exceptional generalization power through exposure to diverse distributions, offering a transformative paradigm for this task. However, our experimental results reveal that LVMs pre-trained on natural-image-dominated data can effectively capture the features of both natural and generated images, yielding comparably low losses and thus limited discriminative capacity between them. This prompts a key question: When and how do LVMs exhibit different behaviors when capturing features of natural and generated images? This investigation reveals an insight: during unlearning, LVMs exhibit disparate forgetting dynamics with feature degradation for generated images escalating faster than natural ones. Inspired by the disparate dynamics, we introduce two detection methods: 1) data-free detection, which prunes model parameters to induce unlearning without data access, and 2) data-driven detection, which optimizes LVMs to unlearn knowledge tied to generated images. Extensive experiments conducted on various benchmarks demonstrate that our unlearning-based approach outperforms conventional detection methods. By recasting the detection task as a problem of machine unlearning, our work establishes a new paradigm for generated image detection.

Comments19 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑