发表机构
Institute of Science and Technology Austria (ISTA)(奥地利科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过将注意力权重映射到边际注意力空间,发现按token缩减时产生跨模型稳健的文本内在信号,按头缩减时形成模型私有签名,并据此提出一种无需训练的KV缓存逐出预算分配方法,性能与重计算或训练预算的方法相当。
AI 中文摘要
虽然独立训练的语言模型之间的表示相似性已被充分记录,但其内部机制(如注意力)在不同模型间的行为却远未得到充分表征。受这一差距的启发,我们通过边际化查询位置,将softmax后的注意力权重映射到一个联合的token-头“边际注意力空间”中,从而考察其结构。在对60多个不同的LLM进行评估时,我们发现当沿token轴和头轴缩减该空间时,会涌现出不同的性质。当按token维度缩减时,边际注意力产生了一种跨模型稳健保守的文本内在信号。为解释这一性质,我们从经验上将边际注意力与网络的输入-输出雅可比矩阵联系起来,并在理论上证明,在平滑性假设下,具有相似下一个token分布的模型必然具有相似的输入-输出雅可比矩阵统计量。当按头维度缩减时,它形成了一种跨文档保守的模型私有签名。在实践中,这为估计每个头的键值(KV)缓存驱逐预算提供了一种自然方法,有效地将模型特定的预算分配与文本内在的token评分解耦。在标准驱逐基准上,基于预训练文本离线预计算的每头预算,结合无需训练的token评分,其性能与在每篇文档上重新计算预算或针对每个目标训练预算的方法相当。代码可在该https URL获取。
英文摘要
While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head "marginal attention space". Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretically that under a smoothness assumption, models with similar next-token distributions are guaranteed to have similar input-output Jacobian statistics. When reduced head-wise, it forms a model-private signature conserved across documents. Practically, this provides a natural way to estimate a per-head budget for key-value (KV) cache eviction, effectively decoupling model-specific budget allocation from text-intrinsic token scoring. On standard eviction benchmarks, a per-head budget precomputed offline on pretraining text, combined with a training-free token score, shows competitive performance with methods that recompute the budget on every document or train it per target. Code available at https://github.com/Flegyas/marginal-attention