arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.00979cs.CVcs.AI

IVY-FAKE:图像和视频AIGC检测的统一可解释框架和基准

IVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection

  • Nanjing University+(南京大学+)

机构由 AI 辅助整理,请以论文原文为准。

Changjiang Jiang, Wenhui Dong, Zhonghao Zhang, Fengchang Yu, Wei Peng, Xinbin Yuan, Yifei Bi, Ming Zhao, Zian Zhou, Chenyang Si, Caifeng Shan

更新

AI总结:

本文提出IVY-FAKE,首个大规模多模态可解释AIGC检测基准,包含106000+丰富标注样本和5000个手动验证示例,通过GRPO强化学习模型实现可解释推理,提升多基准检测性能,显著优于现有方法。

AI中文摘要:

人工智能生成内容(AIGC)技术的快速发展使高质量合成内容的生成成为可能,但同时也引发了重大安全关切。当前检测方法面临两大局限:(1)缺乏多维可解释数据集,现有开源数据集(如WildFake、GenVideo)依赖过于简化的二元标注,限制了训练检测器的可解释性和可信度。(2)先前基于MLLM的伪造检测器(如FakeVLM)在逐步推理中的可解释性不够精细,阻碍了可靠定位和解释。为解决这些挑战,我们引入Ivy-Fake,首个大规模多模态可解释AIGC检测基准。它包含超过106000个丰富标注的训练样本(图像和视频)和5000个手动验证的评估示例,通过精心设计的流程从多个生成模型和真实世界数据集中获取,以确保多样性和质量。此外,我们提出了Ivy-xDetector,基于组相对策略优化(GRPO)的强化学习模型,能够生成可解释的推理链并在多个合成内容检测基准中实现稳健性能。大量实验验证了我们数据集的优越性和方法的有效性。值得注意的是,我们的方法将GenImage的性能从86.88%提升到96.32%,显著超越现有最先进方法。

英文摘要:

The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major limitations: (1) the lack of multidimensional explainable datasets for generated images and videos. Existing open-source datasets (e.g., WildFake, GenVideo) rely on oversimplified binary annotations, which restrict the explainability and trustworthiness of trained detectors. (2) Prior MLLM-based forgery detectors (e.g., FakeVLM) exhibit insufficiently fine-grained interpretability in their step-by-step reasoning, which hinders reliable localization and explanation. To address these challenges, we introduce Ivy-Fake, the first large-scale multimodal benchmark for explainable AIGC detection. It consists of over 106K richly annotated training samples (images and videos) and 5,000 manually verified evaluation examples, sourced from multiple generative models and real world datasets through a carefully designed pipeline to ensure both diversity and quality. Furthermore, we propose Ivy-xDetector, a reinforcement learning model based on Group Relative Policy Optimization (GRPO), capable of producing explainable reasoning chains and achieving robust performance across multiple synthetic content detection benchmarks. Extensive experiments demonstrate the superiority of our dataset and confirm the effectiveness of our approach. Notably, our method improves performance on GenImage from 86.88% to 96.32%, surpassing prior state-of-the-art methods by a clear margin.

补充信息

↑