统一语义先验与高频痕迹:利用混合专家增强V-JEPA实现鲁棒的合成图像取证
Unifying Semantic Priors and High-Frequency Traces: Enhancing V-JEPA with Mixture-of-Experts for Robust Synthetic Image Forensics
查看机构详情
- Sapienza University of Rome(罗马大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对深度伪造检测中全局理解不足的问题,提出MoE-JEPA双流架构,利用JEPA世界模型先验、残差混合专家和噪声流分支,在SID-Set基准上以95.54%准确率超越更大模型。
中文摘要 AI 辅助
社交媒体平台上未经管控的篡改图像泛滥加剧了错误信息的传播,对公众信任和信息完整性构成严重威胁。现代深度伪造检测器通常依赖视觉Transformer(ViT)来捕捉完全合成或局部篡改图像所特有的低级不一致性。然而,此类基础模型对图像的全局理解不足以单独区分真实与伪造的多媒体内容,尤其是在图像经过压缩或通过社交媒体传输等具有挑战性的场景中。本文开创性地将联合嵌入预测架构(JEPA)模型应用于深度伪造检测,利用这类世界模型所展现的对视觉现实的泛化表征能力。我们假设并实证证明,JEPA模型固有的世界理解能力可作为深度伪造检测器的强先验。为充分挖掘JEPA的潜力,我们提出MoE-JEPA,一种用于深度伪造检测的双流架构。通过用残差混合专家(MoE)机制增强V-JEPA 2骨干网络,并引入噪声流分支,我们的模型能够动态内化取证知识。此外,采用门控注意力多实例学习(MIL)模块以确保精确的空间语义理解。在包含30万张AI生成、篡改和真实图像的SID-Set基准上评估,MoE-JEPA以95.54%的准确率确立了新的最先进水平,成功超越了规模大得多的模型。
英文摘要
The unchecked proliferation of manipulated images on social media platforms has increased the spread of misinformation, posing a severe threat to public trust and information integrity. Modern deepfake detectors typically rely on Vision Transformers (ViTs) to capture the low-level inconsistencies that characterize fully synthetic or locally tampered images. However, the global understanding of such foundation models is not enough to discriminate alone between real and fake multimedia content, especially in challenging scenarios where images are compressed or transmitted through social media. In this paper we pioneer the application of Joint-Embedding Predictive Architecture (JEPA) models to deepfake detection, taking advantage of the generalized representation of visual reality that such World Models have exhibited. We hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector. To fully exploit JEPA capabilities, we propose MoE-JEPA, a dual-stream architecture for deepfake detection. By enhancing a V-JEPA 2 backbone with a Residual Mixture-of-Experts (MoE) mechanism, along with a noise stream branch, our model dynamically internalizes forensic knowledge. Furthermore, a Gated Attention Multiple Instance Learning (MIL) module is employed to ensure precise spatial semantic understanding. Evaluated on the SID-Set benchmark, comprising 300K AI-generated, tampered and authentic images, MoE-JEPA establishes a new state-of-the-art with an accuracy of 95.54%, successfully outperforming vastly larger models.