AI 中文总结
该研究针对AI生成图像检测,发现视觉-语言模型PE比DINOv3更具潜力,提出语义原型校准(SPC)方法得到PE-SPC,在多基准测试中达到新的最先进性能。
AI 中文摘要
近期研究表明,对现代视觉基础模型(VFM)的冻结表示进行简单线性探测,可达到最先进的人工智能生成图像(AIGI)检测性能,在具有挑战性的野外场景中大幅优于专用检测器。该发现已确立DINOv3为后续改进的主导基础模型基线。然而,我们发现视觉-语言模型感知编码器(PE)在AIGI检测中具有更大潜力,因为其语言对齐表示保留了高级来源语义。具体而言,PE在其冻结特征空间中表现出比DINOv3更强的局部来源组织。但语义无关的线性探测无法利用该结构,PE-Linear在野外数据集上的表现仍比DINOv3-Linear差4.1%。基于此观察,我们提出语义原型校准(SPC),其从取证语义信息构建类别原型并使用监督数据对其进行校准。我们将SPC应用于PE,得到的检测器称为PE-SPC。我们的分析表明,该简单设计实现了更强的泛化能力。在跨生成器、后处理及野外基准测试中,PE-SPC超越了之前的DINOv3基线,取得了新的最先进结果。
英文摘要
Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in challenging in-the-wild scenarios. This finding has established DINOv3 as the dominant foundation-model baseline for subsequent improvements. However, we find that the vision-language model Perception Encoder (PE) holds greater potential for AIGI detection, because its language-aligned representation preserves high-level provenance semantics. Specifically, PE exhibits stronger local provenance organization than DINOv3 in its frozen feature space. However, semantic-agnostic linear probing fails to exploit this structure, as PE-Linear still underperforms DINOv3-Linear by 4.1% on In-the-Wild. Based on this observation, we propose Semantic Prototype Calibration (SPC), which constructs category prototypes from forensic semantic information and calibrates them with supervised data. We apply SPC to PE and refer to the resulting detector as PE-SPC. Our analysis shows that this simple design achieves stronger generalization. Across cross-generator, post-processing, and in-the-wild benchmarks, PE-SPC surpasses the previous DINOv3 baseline and achieves new state-of-the-art results.