arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24359cs.CVcs.AIcs.CR

解析智能体取证:分流、提示与证据仲裁在开放世界虚假图像检测中的作用

Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection

Xianlong Li, Pietro Bongini, Niccoló Pancino, Marco Blanchini, Benedetta Tondi, Mauro Barni

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过无训练智能体框架剖析分流、提示与推理在开放世界图像取证中的作用,发现推理质量是性能主导因素,核心挑战在于校准信任与仲裁证据而非检测操作。

中文摘要 AI 辅助

图像取证日益成为一个开放世界问题:操作范围从完全合成图像到局部编辑、拼接和换脸,而大多数取证检测器仍专精于单一操作类型。智能体人工智能最近作为一种有前景的解决方案出现。原则上,此类系统可以评估单个检测器的可靠性,识别超出范围的证据,并仲裁相互冲突的报告。然而,目前尚不清楚哪些组件真正驱动性能,以及它们的优势在分布偏移下是否持续。为回答这些问题,我们研究了一个基于专用检测器、逐检测器分流和冲突感知证据仲裁的无训练智能体框架。使用六种配置和三种多模态大语言模型骨干,我们剖析了分流、提示和推理质量在分布内和分布外数据上的作用。我们的结果表明,朴素检测器融合在真实图像上遭受严重的假阳性率。分流和提示通过过滤不可靠证据和暴露检测器局限性,持续提升性能。然而,主导因素是推理本身:更强的评判者显著优于较弱的评判者,尤其是在分布偏移下。最值得注意的是,操作召回率在所有配置中几乎饱和,表明开放世界图像取证的主要挑战不是检测操作,而是校准对专用取证工具的信任并仲裁冲突证据。

英文摘要

Image forensics is increasingly an open-world problem: manipulations range from fully synthetic images to localized edits, splicing and swapping, while most forensic detectors remain specialized to a single manipulation family. Agentic AI has recently emerged as a promising solution. In principle, such systems can assess the reliability of individual detectors, identify out-of-scope evidence, and arbitrate conflicting reports. However, it remains unclear which components actually drive performance and whether their benefits persist under distribution shift. To answer these questions, we study a training-free agentic framework built around specialist detectors, per-detector triage, and conflict-aware evidence arbitration. Using six configurations and three multimodal large language model backbones, we dissect the role of triage, prompting, and reasoning quality on both in-distribution and out-of-distribution data. Our results show that naive detector fusion suffers from severe false-positive rates on authentic images. Triage and prompting consistently improve performance by filtering unreliable evidence and exposing detector limitations. However, the dominant factor is represented by reasoning itself: A stronger judge substantially outperforms a weaker one, particularly under distribution shift. Most notably, manipulation recall is nearly saturated across all configurations, indicating that the main challenge of open-world image forensics is not detecting manipulations, but calibrating trust in specialized forensic tools and arbitrating conflicting evidence.

发表机构

  • University of Siena(锡耶纳大学)
  • IMT School for Advanced Studies Lucca(IMT卢卡高等研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑