发表机构
University of South Dakota; USD Artificial Intelligence Research(南达科他大学; 南达科他大学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文对175篇应用Grad-CAM于视觉Transformer(ViT)的论文开展审计,提出ViT Grad-CAM适配的系统分类,指出其常被误为CNN版Grad-CAM的简单扩展,存在方法选择未明确的问题。
AI 中文摘要
梯度加权类激活映射(Grad-CAM)被广泛用于可视化模型决策,但其最初是针对卷积神经网络(CNN)设计的,在CNN中空间特征图和通道维度具有明确的架构含义。视觉Transformer(ViT)没有相同的结构,而是通过token、注意力机制、残差流和多模态交互来表示图像。本文对Grad-CAM及相关方法如何针对基于ViT的架构进行适配、论证和报告情况开展了系统分类与文献审计。在对超过550篇论文的初步搜索中,我们识别出175篇将Grad-CAM或Grad-CAM邻近方法应用于ViT的论文。我们发现,大多数论文未提供Grad-CAM适配Transformer表示的完整数学或实现层面说明。为表征这一差距,我们引入了ViT Grad-CAM适配的描述性分类,明确了特征位置、梯度目标、空间重建步骤和聚合选择等常被隐含的要素。该分类并非旨在规定单一正确适配方式,而是阐明正在做出的方法选择范围。研究表明,ViT上的Grad-CAM常被视为CNN版Grad-CAM的简单扩展,尽管其需要对严谨性、可复现性和解释性产生影响的非平凡选择。
英文摘要
Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional neural networks, where spatial feature maps and channel dimensions have clear architectural meanings. Vision Transformers (ViTs) do not provide the same structure, instead representing images through tokens, attention, residual streams, and multimodal interactions. This paper presents a systematic taxonomy and literature audit of how Grad-CAM and related methods are adapted, justified, and reported for ViT-based architectures. From an initial search of more than 550 papers, we identify 175 papers that apply Grad-CAM or Grad-CAM-adjacent methods to ViTs. We find that most papers do not provide a full mathematical or implementation-level account of how Grad-CAM is adapted to transformer representations. To characterize this gap, we introduce a descriptive taxonomy of ViT Grad-CAM adaptations that makes explicit the feature locations, gradient targets, spatial reconstruction steps, and aggregation choices that are often left implicit. This taxonomy is not intended to prescribe a single correct adaptation, but to clarify the range of methodological choices being made. The study shows that Grad-CAM on ViTs is often treated as a trivial extension of CNN-based Grad-CAM, despite requiring nontrivial choices that affect rigor, reproducibility, and interpretation.