PE-Mamba:用于AI生成图像检测的双向选择性层聚合
PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection
查看机构详情
- College of Innovation & Technology, University of Michigan-Flint(密歇根大学弗林特分校创新与技术学院)
- School of Electronics and Information Engineering, Korea Aerospace University(韩国航空大学电子与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对AI生成图像检测的挑战,提出基于PE-Core视觉Transformer的PE-Mamba框架,通过双向选择性聚合器等三个组件实现跨层特征聚合,在两个数据集上优于18种检测器且仅训练少量参数。
中文摘要 AI 辅助
由于生成模型的快速发展以及合成内容与真实内容之间的差距不断缩小,AI生成图像(AIGI)检测变得越来越具有挑战性。现基于视觉Transformer的检测器通常依赖加权求和策略来聚合Transformer层间的中间表示,却常常忽略了从浅层纹理线索到深层语义表示的分层特征固有的有序语义进展。在本研究中,我们提出了PE-Mamba,这是一个基于预训练PE-Core视觉Transformer构建的新型框架,采用轻量级LoRA适配,引入了三个互补组件用于跨层特征聚合与融合:第一,双向选择性聚合器(BSA)通过正向和反向选择性扫描处理分层分类标记,其中正向扫描逐步积累从浅到深的取证证据,反向扫描执行从深到浅的上下文细化,以结合高级语义上下文重新解释低级线索;第二,softmax加权聚合器(SWA)计算所有层标记的学习全局摘要,作为互补的聚合路径;第三,sigmoid门控融合器(SGA)通过可学习标量门自适应融合BSA和SWA的输出,使模型能够动态平衡方向序列证据与全局分层聚合。在UniversalFakeDetect(96.6% mACC、99.5% mAP)和AIGCDetect(95.3% mACC、98.1% mAP)上进行的大量实验表明,PE-Mamba在各类生成模型上的泛化能力优于18种检测器,且仅训练总参数的1.3%(仅LoRA占0.13%)。
英文摘要
AI-generated image (AIGI) detection has become increasingly challenging due to the rapid advancement of generative models and the diminishing gap between synthetic and authentic content. Existing vision transformer-based detectors commonly rely on weighted-sum strategies to aggregate intermediate representations across transformer layers, often overlooking the inherently ordered semantic progression of hierarchical features from shallow texture cues to deep semantic representations. In this work, we propose \textbf{PE-Mamba}, a novel framework built upon a pre-trained PE-Core vision transformer with lightweight LoRA adaptation that introduces three complementary components for cross-layer feature aggregation and fusion. First, a bidirectional selective aggregator (BSA) processes layer-wise classification tokens through forward and backward selective scans, where the forward scan progressively accumulates shallow-to-deep forensic evidence, and the backward scan performs deep-to-shallow contextual refinement to reinterpret low-level cues in light of high-level semantic context. Second, a softmax-weighted aggregator (SWA) computes a learned global summary of all layer tokens as a complementary aggregation path. Third, a sigmoid-gated blend (SGA) adaptively fuses the BSA and SWA outputs via a learnable scalar gate, allowing the model to dynamically balance directional sequential evidence and global layer-wise aggregation. Extensive experiments on UniversalFakeDetect (96.6\% mACC, 99.5\% mAP) and AIGCDetect (95.3\% mACC, 98.1\% mAP) demonstrate that \methodname{} outperforms 18 detectors with superior generalization across diverse generative models, while training only 1.3\% of total parameters (0.13\% for LoRA alone).