arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PATE-Forensics:以通用多模态大语言模型(MLLM)为工具的可解释深度伪造取证方法

PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs

Yaqi Li, Jielun Peng, Yabin Wang, Jincheng Liu, Xiaopeng Hong

arXiv 2608.18573首次发表:更新:

发表机构

Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出PATE-Forensics,采用“感知即工具”范式,基于DINOv3构建取证感知工具,结合通用MLLM实现可解释深度伪造取证,在DDL-X Track 3数据集上取得0.89的最佳官方分数,较次席高出0.19分。

AI 中文摘要

现有可解释深度伪造取证方法通常依赖任务适配的多模态大语言模型(MLLM),以同时解决检测、定位和解释问题。受智能体式工具使用的启发,我们提出了“感知即工具”范式,并将其实例化为PATE-Forensics,该方法在架构上将检测与定位同解释生成解耦,同时在取证感知工具内尽可能紧密地耦合检测与定位。基于DINOv3的工具耦合了多粒度检测模块(整合全局、补丁级和片段级证据)与线索引导的定位模块,通过将补丁级和片段级证据空间化为伪造分数图,来引导密集掩码预测。原始图像与该工具生成的取证感知输出构成结构化取证上下文,供通用MLLM使用,在提示约束的引导下,无需任务特定微调即可生成解释。在DDL-X Track 3数据集上,PATE-Forensics取得了官方最佳分数0.89,较排名第二的团队高出0.19分。我们的代码可在指定URL获取。

英文摘要

Existing explainable deepfake forensic methods typically rely on task-adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent-style tool use, we instead introduce a Perception-as-Tool paradigm and instantiate it as PATE-Forensics, which architecturally decouples detection and localization from explanation generation while coupling detection and localization as tightly as possible within a forensic perception tool. The DINOv3-based tool couples a multi-granularity detection module that integrates global, patch-level, and segment-level evidence with a cue-guided localization module by spatializing the patch-level and segment-level evidence into forgery score maps that guide dense mask prediction. The original image and forensic perception outputs produced by the tool form structured forensic context for a general-purpose MLLM, which is guided by prompt constraints to generate explanations without task-specific fine-tuning. On DDL-X Track 3, PATE-Forensics achieves the best official score of 0.89, outperforming the second-ranked team by 0.19 points. Our code is available at https://github.com/yqli00000/PATE-Forensics.

Comments9 pages, 3 figures, 2 tables; DDL-X Track 3, IJCAI 2026 AI Safety Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑