发表机构
Lomonosov Moscow State University; MSU Institute for Artificial Intelligence(莫斯科国立大学; 莫斯科国立大学人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可解释深度伪造检测挑战,提出多骨干检测器结合伪影证据图与对比学习,并利用类条件语言模型生成解释,在 XPlainVerse 上取得高检测与解释分数。
AI 中文摘要
近年来生成式图像模型的进展使得许多被篡改的图像高度逼真,这提高了对不仅准确而且能够为其决策提供视觉证据的检测器的需求。在本文中,我们介绍了我们在 XPlainVerse 数据集 [1] 上的可解释深度伪造检测挑战赛 [2] 中的解决方案,其中要求系统预测图像是真实还是伪造,并生成基于可见取证线索的复杂和简单解释。我们的方法采用模块化的检测与解释设计。对于真实/伪造决策,我们构建了一个多骨干检测器,结合了多个 DINOv3 模型与 Mesorch 操作定位特征,融合了预训练视觉表示、DCT 感知线索和多尺度取证信息。为了将解释证据注入检测器,我们使用基于 Grounding-DINO 的伪掩码生成流程,将训练解释中的局部伪影描述转换为用于伪影证据图的弱补丁级监督。我们进一步引入了一个局部补丁级对比目标,在检测器特征空间中分离伪影和真实性证据,无需配对图像或像素级操作掩码。对于语言输出,我们使用类条件 Qwen3-VL 模型为伪造和真实预测生成复杂解释,随后使用 GRPO 优化的文本简化模型。所提出的方法在 XPlainVerse 的挑战子集上进行了训练和评估。在完整测试分割上,我们的提交达到了 0.9349 的检测准确率、0.5571 的解释分数和 0.7456 的最终挑战分数。
英文摘要
Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.