影响分数与Transformer可解释性:推理时注意力头的有效影响度量
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
浏览论文内容
中文总结 AI 辅助
该研究提出影响分数,结合定向影响与结构贡献分析Transformer注意力头,应用于DeBERTa模型揭示预测行为差异,为研究Transformer分类器决策机制提供系统方法。
中文摘要 AI 辅助
我们提出一种影响分数,用于量化注意力头对基于Transformer的提示注入检测模型分类决策的贡献。该分数结合了对logits的定向影响与残差流内的结构贡献,支持在头、层及网络层级进行多尺度分析。将其应用于专门用于提示注入检测的DeBERTa模型后,我们的框架揭示了正确与错误预测间的不同决策行为。该方法在细粒度电路分析与基于全局输出的方法间提供了有效折中,为研究Transformer分类器的决策机制提供了系统途径。
英文摘要
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.
发表机构
- Thales CDI(泰雷兹CDI公司)
- Inria Paris(法国国家信息与自动化研究所巴黎中心)
- Sorbonne Université(索邦大学)
机构由 AI 辅助整理,请以论文原文为准。