arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPARED:基于推理的AI生成图像检测方法,采用对抗编辑数据

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan

arXiv 2608.12876首次发表:更新:

AI 中文总结

本研究提出对抗强化学习框架SPARED,通过扩散图像编辑器与推理型MLLM的交替博弈,训练出能抗捷径、泛化能力强的AI生成图像检测器,在三个外部基准上性能单调提升。

AI 中文摘要

检测AI生成图像仅为任务的一半:部署的检测器还必须为其判断提供依据,但现有检测器从训练数据中继承了三种失效模式:来自不同来源的真实图像和伪造图像会引发溯源捷径,监督式解释语料库会传授模板化理由,而静态伪造语料库会使决策边界保持静止,与此同时生成器却在不断演进。我们引入SPARED,一种对抗强化学习框架,该框架让两个异构模型相互博弈。扩散图像编辑器学习将真实照片编辑为能欺骗当前检测器的对应伪造图像,而推理型多模态大语言模型(MLLM)学习以基于自由形式推理的判断来揭露这些伪造图像。两种奖励机制从设计上均为抗捷径型:仅当编辑被忠实地执行时,攻击者(编辑器)才获得奖励;仅当判断正确时,防御者(MLLM)才获得奖励。随着两个模型交替博弈,每一轮的攻击者都会生成更难的训练池,以针对当前检测器的盲区,因此检测器必须泛化而非记忆任何固定的人工制品分布。尽管解释从未被作为奖励目标,但其质量却随着仅以准确性为目标的训练轮次增加而逐步提升。在三轮博弈中训练的检测器,在三个外部基准上的性能均呈单调提升。

英文摘要

Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving. We introduce \methodname{}, an adversarial reinforcement learning framework that pits two heterogeneous models against each other. A diffusion image editor learns to edit real photographs into fake counterparts of those same photographs that fool the current detector, while a reasoning MLLM learns to expose them with a verdict grounded in free-form reasoning. Both rewards are shortcut-proof by design: the attacker is credited only when its edit is faithfully executed, and the defender only when its verdict is correct. As the two models alternate, each round's attacker regenerates a harder training pool aimed at the current detector's blind spots, so the detector must generalize rather than memorize any fixed artifact distribution. Although the explanation is never rewarded, its quality rises round over round as a side effect of accuracy-only training. A detector trained within this loop improves monotonically across rounds on each of three external benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑