arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05074cs.CL

影响分数与Transformer可解释性:推理时注意力头的有效影响度量

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出影响分数,结合定向影响与结构贡献分析Transformer注意力头,应用于DeBERTa模型揭示预测行为差异,为研究Transformer分类器决策机制提供系统方法。

中文摘要 AI 辅助

我们提出一种影响分数,用于量化注意力头对基于Transformer的提示注入检测模型分类决策的贡献。该分数结合了对logits的定向影响与残差流内的结构贡献,支持在头、层及网络层级进行多尺度分析。将其应用于专门用于提示注入检测的DeBERTa模型后,我们的框架揭示了正确与错误预测间的不同决策行为。该方法在细粒度电路分析与基于全局输出的方法间提供了有效折中,为研究Transformer分类器的决策机制提供了系统途径。

英文摘要

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.

发表机构

  • Thales CDI(泰雷兹CDI公司)
  • Inria Paris(法国国家信息与自动化研究所巴黎中心)
  • Sorbonne Université(索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑