AI图像检测的可解释性:热图实际显示了什么
Explaining AI-Image Detection: What the Heatmap Actually Shows
浏览论文内容
中文总结 AI 辅助
该研究针对AI图像检测的可解释性问题,构建检测器并测试归因图的可靠性,发现现有归因图无法通过检测器盲对照,未找到可靠的AI图像检测解释方法。
中文摘要 AI 辅助
商品评价照片是一种凭证,平台会依据它批准退款,而生成式模型将伪造这类照片的成本降至为零。我们研究该检测问题,因此构建了一个检测器并附上归因图作为其证据,随后在旨在在出错时改变我们结论的控制条件下,对186527张图像测量这一组合的效果。压缩历史而非合成驱动了朴素评估:我们最强的模型在产品不重叠的划分上达到0.9999 PR-AUC(精确率-召回率曲线下面积),但当我们将合成图像重新编码为真实类别的格式后,该模型的PR-AUC降至0.7254,而五个公开检测器的变化最多仅为0.07。对齐一个类别会转移线索而非移除它,修复后的模型将原生文件的合成概率中位数设为0.0004。对两个类别采用相同的最终编码修复了这一问题,三种子集的析因分析表明,编码变化带来了全部增益(+0.176 ± 0.009 PR-AUC)。该编码仅均衡了最后阶段:仅取证特征仍能以0.7145的性能区分两个类别,而基础率为0.254。为验证证据,我们对归因图进行因果测试,对照条件从不参考检测器。是否存在归因排序完全取决于检测器是否对图像做出反应。在我们的首个修复检测器(100张编辑帧中96张判定为真实)上,没有任何归因图优于随机水平。在我们选定的检测器上,17张编辑图像中有12张的归因图通过了该控制测试,生成图像中有8张通过;扰动在两个轴上均有效,且无任何梯度-CAM变体表现出优势。简单对照条件从未通过测试,在生成图像上中心先验的表现比随机更差。我们的集成区域归因图通过了两个轴的测试,其最高像素AP为每张图12.4秒,而遮挡法为44.9秒。通过检测器盲对照测试目前仍不算是可靠的解释,且我们未找到任何符合要求的解释。
英文摘要
A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study that detection problem, so we build a detector and attach an attribution map as its evidence, then measure what that pair delivers on 186,527 images under controls designed to change our conclusions when something is wrong. Compression history, not synthesis, drives naive evaluation: our strongest model reaches 0.9999 PR-AUC (area under the precision-recall curve) on a product-disjoint split, yet falls to 0.7254 once we re-encode synthetics into the real class's format, while five public detectors move by at most 0.07. Aligning one class relocates the cue rather than removing it, and the repaired model then assigns native files a median probability of synthesis of 0.0004. One identical final encode for both classes repairs that, and a three-seed factorial credits the encoding change with the whole gain (+0.176 +- 0.009 PR-AUC). That encode equalises the last stage only: forensic features alone still separate the classes at 0.7145 against a base rate of 0.254. For evidence we test maps causally, against controls that never consult the detector. Whether an attribution ranking exists at all depends on whether the detector reacts to the image. On our first-fix detector, which calls 96 of 100 edited frames real, no map beats a random one. On the detector we selected, twelve of seventeen maps clear that control on edited images and eight on generated ones; perturbation leads both axes and no gradient-CAM variant shows a positive advantage. The trivial controls never clear it, and on generated images the centre prior is worse than random. Our ensembled regional map clears both axes and takes the top pixel AP at 12.4 s per map against 44.9 for occlusion. Clearing a detector-blind control is not yet a faithful explanation, and we demonstrate none.
发表机构
- Sirius Educational Centre(天狼星教育中心)
- HSE University(高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。