发表机构
Ruhr University Bochum(鲁尔大学波鸿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对基于自编码器重构误差的AI生成图像检测器提出两种新型对抗攻击方法,发现其易受难以察觉的对抗样本攻击,检测性能显著下降,且对抗样本可跨检测器迁移,揭示了此类检测器的固有脆弱性。
AI 中文摘要
AI生成图像的出色视觉质量与广泛普及催生了对可靠且鲁棒的检测方法的需求。基于重构的检测器作为一种有前景的方向出现,可实现透明且无训练的合成图像识别。然而,与标准的基于分类器的方法相比,其运作模式存在根本差异,目前人们对其对抗鲁棒性知之甚少。本研究针对利用自编码器重构误差的检测器提出两种新型攻击方法。我们发现,通过构造难以察觉的对抗样本,可人为增大原始图像与重构图像之间的距离,致使伪造图像被错误分类为真实图像。我们的评估涵盖来自三个最先进生成器的图像与三个检测器,结果表明,即便被攻击图像还经历了现实世界的退化,检测性能仍显著下降。关键的是,我们的对抗样本可自然跨检测器迁移,因为所有检测器都遵循相同的原理,这指向基于重构的检测器存在固有脆弱性。
英文摘要
The impressive visual quality and ubiquity of AI-generated images call for reliable and robust detection methods. Reconstruction-based detectors have emerged as a promising direction for transparent and training-free identification of synthetic images. However, due to their fundamentally different mode of operation (compared to standard, classifier-based methods), little is known about their adversarial robustness. In this work, we propose two novel attack methods targeted at detectors that leverage autoencoder reconstruction error. We find that by constructing imperceptible adversarial examples, the distance between original and reconstruction can be artificially increased, causing fake images to be wrongly classified as real. Our evaluation including images from three state-of-the-art generators and three detectors demonstrates that detection performance is significantly decreased, even if attacked images additionally undergo real-world degradations. Critically, our adversarial examples naturally transfer across detectors, as they all share the same principle, pointing towards an inherent vulnerability of reconstruction-based detectors.
CommentsAccepted at AI4MFDD (AI for Multimedia Forensics & Disinformation Detection) Workshop, ECCV 2026