智能体工具增强推理用于可解释的图像伪造检测
Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection
浏览论文内容
中文总结 AI 辅助
提出ATAR框架,集成22种取证工具,通过双流推理和课程学习,在零样本IMDL基准上以78.5%平均F1超越最强MLLM基线11.8个百分点,实现可解释的图像伪造检测。
中文摘要 AI 辅助
传统的图像伪造检测方法产生二值分数或像素级掩码,缺乏可解释的证据,而近期基于多模态大语言模型(MLLM)的方法则对预定的分类结果生成事后解释,而非从证据出发进行推理。受人类司法专家取证工作流程的启发,我们提出了智能体工具增强推理(ATAR),这是一个集成七个互补领域中22种专业取证工具的框架,通过多轮推理自主检测、定位和解释图像伪造。双流取证推理范式结合了高层语义异常路径(放大可疑区域以进行细粒度检查)和低层伪造伪影路径(调用取证工具提取客观证据)。我们进一步引入了取证课程学习:在通用经验监督微调阶段,自动化的师生指导流程合成多轮工具使用推理轨迹;在取证场景强化学习阶段,工具先验课程引导早期工具探索并逐步将控制权转移给智能体,同时结构化证据奖励提供细粒度的过程级监督。在IMDL、深度伪造检测、DMDL和AIGC检测上的实验表明,ATAR在六个零样本IMDL基准上实现了78.5%的平均图像级F1分数,超过最强MLLM基线11.8个百分点,并在其他任务上与专业检测器保持竞争力,同时产生更忠实和有依据的解释。
英文摘要
Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.
发表机构
- Nanyang Technological University(南洋理工大学)
- Ant Digital Technologies, Ant Group(蚂蚁数字科技,蚂蚁集团)
- Singapore University of Technology and Design(新加坡科技设计大学)
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。