发表机构
Ontario Tech University(安大略理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对医学与航空影像中旋转诱导的显著性漂移问题,提出EquiGrad-CAM方法提升显著性图等变性,效果优于旋转增强训练,副产品PEUM可评估解释可复现性。
AI 中文摘要
事后显著性图(如Grad-CAM)越来越多地用于审计已部署视觉模型做出决策的原因,但即使预测结果不变,当输入旋转时,热力图也会发生漂移。在组织病理学和航空影像等没有规范方向的领域,这破坏了将显著性作为证据的使用。本文研究该漂移是忠实信号还是CAM算子引入的噪声,通过测量算子每个阶段的等变性而非从网络输出推断来回答这一问题。不稳定性出现在意外之处:通道权重是最具旋转稳定性的阶段,在ResNet-50上完全稳定,因为全局平均池化(GAP)+线性头使类别梯度场在空间上恒定;发生移动的是空间激活张量,而分类器自身的池化会丢弃该移动。一项因果测试证实了这一结果:在任意方向上,遮挡显著性漂移的像素比遮挡随机像素对模型造成的损失更小。该漂移由分类器丢弃的自由度承载,这使得去除漂移是忠实的而非破坏性的。本文提出EquiGrad-CAM,这是一种无需训练的包装器,它取T个旋转视图,将每个视图的显著性逆旋转到共同规范帧后取平均。在完整ImageNet-1K验证集上,它使单视图Grad-CAM的等变性提升了36.0%(ResNet-50)、87.5%(VGG-16)和247%(ViT-B/16);规模匹配的消融实验表明,平均前的对齐而非聚合位置是驱动因素。它无需重新训练即可优于旋转增强训练,将零样本CLIP提升了145%,并在PatchCamelyon和RESISC45上产生旋转一致的解释。其副产品PEUM可对解释的可复现性进行排名,成本仅为已获取的视图之外的额外开销。代码:this https URL
英文摘要
Post-hoc saliency maps such as Grad-CAM are increasingly used to audit why a deployed vision model made a decision, yet the heatmap drifts when the input is rotated, even when the prediction is unchanged. In domains with no canonical orientation, such as histopathology and aerial imagery, this undermines using saliency as evidence. We ask whether that drift is faithful signal or noise introduced by the CAM operator, and answer it by measuring equivariance at every stage of the operator rather than inferring it from the network's output. The instability is not where one would guess: the channel weights are the most rotation-stable stage, and on ResNet-50 exactly stable, because a GAP+linear head makes the class gradient field spatially constant. What moves is the spatial activation tensor, and the classifier's own pooling discards that movement. A causal test confirms the consequence: occluding the pixels whose saliency drifts costs the model less than occluding random pixels, at either orientation. The drift is carried by degrees of freedom the classifier throws away, which is what makes removing it faithful rather than destructive. EquiGrad-CAM is a training-free wrapper that takes T rotated views, inverse-rotates each view's saliency into a common canonical frame, and averages. On the full ImageNet-1K validation set it raises equivariance over single-view Grad-CAM by +36.0% (ResNet-50), +87.5% (VGG-16) and +247% (ViT-B/16); a scale-matched ablation isolates alignment before averaging, not the locus of aggregation, as the driver. It beats rotation-augmented training without retraining, lifts zero-shot CLIP by +145%, and yields rotation-consistent explanations on PatchCamelyon and RESISC45. Its by-product PEUM ranks explanations by how reproducible they are, at no cost beyond the views already taken. Code: https://github.com/Khawaja-Murad/EquiGrad-CAM
Comments11 pages, 3 figures, 6 tables. Code at https://github.com/Khawaja-Murad/EquiGrad-CAM