ResLRP:残差消除在视觉Transformer归因不稳定性中的作用
ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers
浏览论文内容
中文总结 AI 辅助
ResLRP通过显式处理残差连接中的消除效应,解决了视觉Transformer中LRP归因不稳定的问题,显著提升忠实度和定位性能,尤其在视觉语言模型中效果最佳。
中文摘要 AI 辅助
视觉Transformer(ViT)是现代大多数视觉模型的核心,但获得细粒度、忠实且稳定的输入归因仍然具有挑战性。逐层相关性传播(LRP)已被适配到Transformer注意力机制中,但在ViT中,它常常产生嘈杂且不忠实的解释。我们表明,缺失的关键在于对残差连接的处理:残差路径中的消除效应导致归因爆炸。此外,我们发现这些消除效应在ViT中比在语言Transformer中显著更强。为解决此问题,我们引入了残差感知的逐层相关性传播(ResLRP),这是LRP的一个简单扩展,其传播规则明确考虑了残差分支中的消除效应,严格保持保守性,并可证明地限制相关性爆炸。因果通道级干预证实,是残差消除(而非一般的正则化效应)驱动了不稳定性。ResLRP在忠实度和定位方面显著提高了归因质量,在涵盖监督、自监督、对比、层次和多模态家族的ViT架构上,以及在真值控制的FunnyBirds基准上进行了评估。最大的提升出现在现代视觉语言模型(VLM)中,定位提升+27-29%,忠实度分数提升高达3.4倍。在基准之外,ResLRP在输入空间中定位稀疏自编码器(SAE)特征,我们的残差放大度量可作为架构级诊断,预测归因退化的位置。
英文摘要
Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Relevance Propagation (LRP) has been adapted to transformer attention, but in ViTs it often produces noisy, unfaithful explanations. We show that the missing ingredient is the treatment of residual connections: cancellation effects in residual pathways lead to attribution explosion. Moreover, we find that these cancellations are substantially stronger in ViTs than in language transformers. To address this issue, we introduce Residual-aware Layer-wise Relevance Propagation (ResLRP), a simple extension of LRP whose propagation rules explicitly account for cancellations in residual branches, are exactly conservative, and provably bound relevance explosion. Causal channel-wise interventions confirm that residual cancellation, not a generic regularization effect, drives the instability. ResLRP substantially improves attribution quality across faithfulness and localization, evaluated on ViT architectures spanning supervised, self-supervised, contrastive, hierarchical, and multimodal families, as well as on the ground-truth-controlled FunnyBirds benchmark. The largest gains arise in modern Vision Language Models (VLMs), with +27-29% localization and up to 3.4x faithfulness scores. Beyond benchmarks, ResLRP localizes Sparse Autoencoder (SAE) features in input space, and our residual amplification measure serves as an architecture-level diagnostic predicting where attribution degrades.
发表机构
- Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所)
- DSC ScaDS.AI, Leipzig University(莱比锡大学DSC ScaDS.AI)
- ICT Cluster, Singapore Institute of Technology(新加坡理工大学ICT集群)
- Institute for Cancer Genetics and Informatics (ICGI), Oslo, Norway(奥斯陆癌症遗传学与信息学研究所)
- Technische Universität Berlin(柏林工业大学)
- BIFOLD – Berlin Institute for the Foundations of Learning and Data(BIFOLD – 柏林学习与数据基础研究所)
机构由 AI 辅助整理,请以论文原文为准。