HiLRP:为视觉Transformer提供一种可信的解释——基于注意力原语的守恒归因
HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives
- Faculty of Engineering, University of Peradeniya(佩拉德尼亚大学工程学院)
- School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院)
- School of Engineering & Digital Technologies, University of Southern Queensland(南昆士兰大学工程与数字技术学院)
- School of Engineering & Technologies, UNSW(新南威尔士大学工程与技术学院)
- Faculty of Science and Technology, Charles Darwin University(查尔斯达尔文大学科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HiLRP是一种统一归因框架,通过分解ViT算子为四类操作并制定守恒规则,解决现有方法在多样化ViT上不可靠的问题,在EfficientViT上指向性达0.97,优于对比方法。
AI中文摘要:
视觉Transformer(ViT)的设计日益多样化,主干网络以各种配置结合卷积茎、窗口注意力、线性注意力或多轴注意力、patch合并以及空间缩减。这种多样性对现有归因方法构成挑战,这些方法的假设往往不适用于不同ViT变体:Grad-CAM需要终端空间特征图,注意力rollout假设全局softmax注意力,层相关传播(LRP)需要特定模块规则。据我们所知,现有方法均未提供覆盖该架构空间的统一归因框架。我们表明,这种架构多样性可由更简单的底层结构捕获。当前ViT中的注意力和分辨率缩减算子可分解为四类操作:线性映射、双线性混合、归一化或门控、以及重索引。每个操作都有满足守恒性的相关规则。基于这些规则,HiLRP通过构造而非特定架构推导支持新主干网络,其归因图分解预测而非依赖启发式假设。我们证明了守恒性和条件等变性,并以机器精度验证两者。在14种归因方法和10种架构上,我们发现没有现有方法能在ViT系列间保持可靠,而Faithfulness Correlation对空间掩码鲁棒的主干网络变得无信息。仅HiLRP在窗口注意力、空间缩减、多轴和线性注意力模型间保持守恒,而朴素扩展会产生零或膨胀的相关性。它还定位了类激活映射中的归因失败,在EfficientViT上达到0.97的指向性,而对比方法仅为0.55。
英文摘要:
Vision Transformer (ViT) design has become increasingly diverse, with backbones combining convolutional stems, windowed, linear, or multi-axis attention, patch merging, and spatial reduction in various configurations. This diversity poses challenges for existing attribution methods, whose assumptions often do not hold across ViT variants: Grad-CAM requires a terminal spatial feature map, attention rollout assumes global softmax attention, and layer-wise relevance propagation (LRP) requires module-specific rules. To the best of our knowledge, no existing method provides a unified attribution framework across this architectural space. We show that this architectural diversity can be captured by a simpler underlying structure. The attention and resolution-reduction operators in current ViTs can be decomposed into four operation types: linear maps, bilinear mixing, normalization or gating, and reindexing. Each operation admits a relevance rule that satisfies conservation. Based on these rules, HiLRP supports new backbones by construction rather than by architecture-specific derivation, and its attribution maps decompose the prediction rather than relying on heuristic assumptions. We prove conservation and conditional equivariance and verify both to machine precision. Across 14 attribution methods and 10 architectures, we find that no prior method remains reliable across ViT families, while Faithfulness Correlation becomes uninformative for backbones robust to spatial masking. HiLRP alone preserves conservation across windowed, spatial-reduction, multi-axis, and linear-attention models, where naive extensions can produce zero or inflated relevance. It also localizes attribution failures in class activation mapping, achieving 0.97 Pointing compared with 0.55 for competing methods on EfficientViT.