Defake-o3:从推测性理由到可验证证据的可解释AIGI检测
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
浏览论文内容
中文总结 AI 辅助
Defake-o3是结合交互式视觉搜索与证据验证器的可解释AIGI检测器,构建了GroundFake数据集和FakeFrontier基准,在多基准上同时提升了AIGI检测准确率与解释质量。
中文摘要 AI 辅助
图像生成模型的快速发展催生了对AI生成图像(AIGI)检测器的需求,这类检测器不仅需要准确,还需具备可解释性和可靠性。尽管基于多模态大语言模型(MLLM)的检测器能提供自然语言解释,但现有方法常生成推测性理由:它们依赖模糊或虚构的瑕疵,遗漏最新生成器产生的细微局部缺陷,且无法提供可视觉验证的证据。本文提出Defake-o3,一种可解释AIGI检测器,实现从推测性理由向可验证证据的转变。它将交互式视觉搜索与验证器引导的证据对齐相结合:模型迭代放大可疑区域以检查细粒度细节,而基于人工验证标注训练的证据验证器(Evidence Verifier)提供强化学习奖励,奖励基于有依据的证据,惩罚无根据的主张。为支撑该目标,我们构建了GroundFake数据集,用于有依据的可解释检测,包含局部边界框证据、基于视觉定位和瑕疵特异性的人工验证、修正后的推理轨迹,以及有效/无效证据监督。我们还引入了FakeFrontier,一个由真实图像和10个最新生成器的输出构建的分布外基准,以及用于评估证据质量和说服力的基于MLLM的协议。在GroundFake、FakeFrontier及其他分布外基准上的实验表明,Defake-o3在提升检测准确率的同时,也提高了解释质量,生成更具局部性、可验证性和说服力的证据。
英文摘要
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。