arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力退化、功能令牌锚定以及基于注意力的大语言模型干预的局限性

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

Sagar Dangal, Manoj Shakya

arXiv 2607.20524首次发表:更新:

发表机构

London Metropolitan University; Islington College; Kathmandu University(伦敦城市大学; 伊斯灵顿学院; 加德满都大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨大语言模型中注意力退化问题,通过六个实验刻画其模式及功能令牌锚定特点,测试干预机制效果,发现平均注意力退化多为描述性,功能令牌作用于隐藏状态计算,为可解释性方法及推理优化提供启示。

AI 中文摘要

在Transformer可解释性中,平均跨位置注意力退化被广泛报道,但它是否因果性地限制上下文检索仍未得到检验。我们在GPT-2、LLaMA-3.2-1B/3B、OPT-1.3B和distilgpt2上进行了六个协同实验。首先刻画短期(5 - 100个令牌)注意力退化,发现通用的指数然后平稳模式,其速率与深度负相关,各架构有不同的层熵特征。功能令牌锚定与架构相关。在子句边界插入逗号可因果性减少40 - 80令牌范围内的预测退化。通过Relay-Aware Attention(RAA)测试因果机制,结果表明其虽增加注意力质量但对不同模型效果各异。多因素检索探针显示退化率不能预测跨模型的检索准确性。我们得出结论,平均注意力退化很大程度上是描述性而非规定性的,功能令牌通过其隐藏状态计算而非所接收的注意力起作用,这对可解释性方法和基于注意力分数的推理优化有影响。

英文摘要

Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested. We present six coordinated experiments across GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, and distilgpt2. We first characterise short-term (5-100 token) attention degradation, finding a universal exponential-then-plateau pattern whose rate is inversely correlated with depth, with distinct layer-wise entropy signatures per architecture. Function token anchoring proves architecture-dependent: OPT-1.3B (absolute positional encoding) shows distance-dependent preposition specificity, GPT-2 shows uniform non-specific dependence, and LLaMA (RoPE) shows reversal at long distances. Strategic comma insertion at clause boundaries causally reduces prediction degradation in the 40-80 token range, with the benefit tied to syntactic boundary alignment rather than token density. We then test the mechanism causally: Relay-Aware Attention (RAA), which biases attention logits toward function token positions, verifiably increases attention mass by 16-24% yet yields null effects on GPT-2 and LLaMA-1B, preliminary harm on LLaMA-3B, and a mixed effect on OPT-1.3B that nets to approximately zero. Multi-fact retrieval probes further show that degradation rate does not predict retrieval accuracy across models. We conclude that mean attention degradation is largely descriptive rather than prescriptive: function tokens contribute through what their hidden states compute, not through the attention they receive -- with implications for interpretability methodology and attention-score-based inference optimisations such as KV-cache eviction.

Comments19 pages, 2 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑