发表机构
University of Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences; Ocean University of China(中国科学院大学; 中国科学院计算技术研究所; 中国海洋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DBCF框架,融合CLIP全局语义与DINOv3局部结构特征,通过双分支互补及参数高效融合模块,提升深度伪造检测的跨数据集和跨篡改泛化能力。
AI 中文摘要
随着图像生成和编辑技术的显著进步,面部伪造对隐私和公共安全构成了重大挑战。由于捕捉伪造线索的能力有限,现有小规模伪造检测模型通常难以在各类领域和未见过的篡改操作中实现泛化。为解决这一局限性,研究人员转向大规模基础模型,这些模型能够提供更丰富的表征和更好的泛化能力。然而,仅依赖单一基础模型仍不足以进行有效的伪造检测。虽然像CLIP这样的模型提供了强大的全局语义线索,但它们缺乏捕捉详细局部面部特征的能力。相比之下,DINO擅长捕捉面部的局部结构特征,但提供的全局语义上下文较弱。为充分利用多个基础模型之间的协同效应,我们提出了一种分层多粒度框架,该框架整合了互补的预训练表征。具体而言,基于CLIP的全局上下文分支(GCB)捕捉整体语义线索,而基于DINOv3的细粒度线索分支(FCB)捕捉局部结构异常。此外,我们设计了一个特征融合模块,通过自适应提取和整合来自两个模型的互补特征,实现对冻结基础骨干网络的参数高效适配。通过联合利用全局上下文和细粒度线索,我们的方法学习了更全面的伪造表征,并实现了强大的跨篡改性能。在多个基准上的大量实验证明了所提出设计的优势,特别是在跨数据集和跨篡改设置下。
英文摘要
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.