AI 中文总结
本综述以方法为中心,整合57篇文献,梳理CAM类视觉解释方法的发展,构建分类体系,分析其从CNN到基础模型时代的演进趋势,同时指出评估碎片化的问题。
AI 中文摘要
类别激活映射(Class Activation Mapping, CAM)是可解释人工智能中应用最广泛的视觉解释方法族之一,其核心目标直观:将模型内部证据转化为热力图,突出支持目标类别或概念的图像区域、卷积通道、token或图像块。自2016年首个CAM公式提出以来,该领域已远超全局平均池化CNN分类器范畴,CAM类方法现涵盖基于梯度的事后解释、无梯度的评分与消融方法、高分辨率上采样、弱监督定位与分割、Transformer token归因、因果与去偏方法,以及使用CLIP、DINO、SAM或特征分布比较的基础模型时代方法。本综述整合了2016年以来发表的57篇以方法为中心的严格论文,构建了按归因机制、架构依赖性和评估目标划分的分类体系,随后对基于梯度的CAM、近期及混合CAM类方法、基于模型或架构感知的方法进行了综述。纵观所有文献,核心趋势清晰:该领域正从解释单个低分辨率CNN层中的单个类别分数,转向比较式、多层、概率、token感知及基础模型感知的解释;与此同时,评估仍呈碎片化状态,忠实性、定位、鲁棒性、计算成本及人类信任度常采用不同协议测量。因此,本综述不仅强调每种方法的贡献,还指出其留下的缺口及后续方法试图填补该缺口的方向。
英文摘要
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.