用于视觉Transformer高效可解释性的梯度跳跃相关性传播
Gradient-Skipping Relevance Propagation for Efficient Explainability of Vision Transformers
浏览论文内容
中文总结 AI 辅助
研究针对视觉Transformer难以解释的问题,提出基于自适应头加权和跳跃感知传播的GradSkip方法,通过对注意力头不同重要性建模及动态分配相关性,在实验中展现出最优忠实度且计算量大幅减少,提升了定位及与真实区域的对齐。
中文摘要 AI 辅助
视觉Transformer(ViTs)难以解释,因为当前的相关性传播和注意力流方法未充分考虑一些关键架构特征,如注意力头和残差连接的重要性不均。现有方法通常假设注意力头重要性一致,且将跳跃连接建模为恒等路径,导致相关性归因不准确。为解决这些问题,我们引入了GradSkip,一种基于自适应头加权和跳跃感知传播的ViTs新型相关性传播方法。GradSkip对注意力头的不同重要性进行建模,并在注意力和残差路径之间动态分配相关性。在ImageNet1K和BloodMNIST上的实验表明,GradSkip具有最优的忠实度,同时所需的GFLOP比现有最佳方法少14倍以上。使用基于Transformer的分割进行的额外评估证实了定位的改善以及与真实区域的对齐。
英文摘要
Vision Transformers (ViTs) are difficult to interpret because current methods of relevance propagation and attention flow do not fully consider some key architectural features, such as the uneven importance of attention heads and residual connections. Prior approaches typically assume uniform importance across attention heads; furthermore, they model skip connections as identity paths, leading to inaccurate relevance attribution. To address these issues, we introduce GradSkip, a novel relevance propagation method for ViTs based on adaptive head weighting and skip-aware propagation. GradSkip models the different importance of the attention heads and dynamically distributes relevance between the attention and residual paths. Experiments on ImageNet1K and BloodMNIST demonstrate a state-of-the-art faithfulness of GradSkip while requiring over 14 times fewer GFLOPs than the best-performing existing approaches. Additional evaluations using transformer-based segmentation confirm improved localization and alignment with ground-truth regions.
发表机构
- DII Polytechnic University of Marche(马尔凯理工大学工程学院)
- CHIMOMO University of Modena and Reggio Emilia(摩德纳与雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。