发表机构
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences (UCAS); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS); Ant Group; Sangfor Technologies Inc.(中国科学院大学先进交叉科学学院; 中国科学院自动化研究所; 中国科学院大学人工智能学院; 蚂蚁集团; 深信服科技股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有AIGI检测模型的感知瓶颈,提出Veritas++框架,通过PoRL与VaOPD机制增强感知能力,提升了检测泛化性与效率。
AI 中文摘要
图像生成模型的能力不断提升,合成图像已成为开放媒体中的常规存在,这使得鲁棒且可泛化的AI生成图像(AIGI)检测愈发重要。多模态大语言模型(MLLMs)为黑盒二元评分提供了透明的替代方案,但我们发现当前基于MLLM的检测器在捕获细粒度异常方面仍存在显著的感知瓶颈,它们主要关注视觉证据的组织与合成方式,而未对内在感知进行优化。为缓解这一差距,我们提出Veritas++,这是一个以可靠感知作为真实性推理基础的感知增强推理框架。我们未直接优化模型的解释能力,而是将AIGI检测建立在三种基本感知能力之上,即捕获细粒度视觉细节、语义异常和像素级差异。基于这一见解,我们引入感知导向学习(PoRL),它用可验证的奖励替代开放式描述监督,以明确强化这些能力。为进一步将增强的感知与推理相结合,我们引入值感知在线策略蒸馏(VaOPD),这是一种自适应蒸馏机制,它优先考虑高价值蒸馏信号而非统一监督,通过特权自教师内化感知感知的推理。在标准、野外和新兴基准上的大量实验表明,Veritas++实现了良好的泛化性,感知学习有效弥合了感知差距并在检测上产生了无缝增益,而VaOPD进一步实现了高效的能力演进且不牺牲现有性能。代码和检查点可在this https URL获取。
英文摘要
The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection increasingly essential. While multi-modal large language models (MLLMs) offer a transparent alternative to black-box binary scoring, we observe that current MLLM-based detectors still exhibit notable perception bottlenecks in capturing fine-grained anomalies. They primarily focus on how visual evidence is organized and synthesized, leaving the intrinsic perception less optimized. To mitigate this gap, we present Veritas++, a perception-enhanced reasoning framework that establishes reliable perception as the foundation of authenticity reasoning. Rather than directly optimizing the model's explanatory ability, we ground AIGI detection on three basic perception abilities, i.e., capturing fine-grained visual details, semantic anomalies and pixel-level differences. Building on this insight, we introduce Perception-oriented Learning (PoRL), which replaces open-ended description supervision with verifiable rewards to explicitly strengthen these capacities. To further integrate enhanced perception with reasoning, we introduce Value-aware On-Policy Distillation (VaOPD), an adaptive distillation mechanism that prioritizes high-value distillation signals over uniform supervision, internalizing perception-aware reasoning through a privileged self-teacher. Extensive experiments across standard, in-the-wild and emerging benchmarks demonstrate that Veritas++ achieves promising generalization. The perception learning effectively bridges the perception gap and yields seamless gains on detection, while VaOPD further enables efficient capability evolvement without sacrificing existing performance. Code and checkpoints are available at https://github.com/EricTan7/VeritasPP.