arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自注意力在竞争条件下的检索容量

Retrieval Capacity of Self-Attention Under Competition

Timur Mudarisov, Mikhail Burtsev, Radu State

arXiv 2609.37879首次发表:更新:

发表机构

University of Luxembourg; London Institute for Mathematical Sciences(卢森堡大学; 伦敦数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过自注意力机制测量语言模型的有效注意力集合大小,发现保留少量高注意力标记即可保持性能,且集合大小受上下文长度、竞争和权重归一化影响,为理解模型检索容量提供新方法。

AI 中文摘要

语言模型在其上下文中实际使用了多少个标记,什么决定了这个数量?我们通过自注意力来研究这个问题。在不进行重新训练的情况下,我们仅保留每个头、层和查询中注意力权重最高的标记,并保持其原始权重不变。通过改变所选集合的大小并测量负对数似然(NLL)的增加,我们估计了在选定损失容差内保持所需的有效注意力集合大小。相对较小的所选集合可以使NLL接近全注意力基线,尽管所需大小因模型而异。基于注意力的选择显著优于随机选择。所选集合表现出几何结构,尽管几何分离本身并不能确定模型损失得以保持。在评估相同预测目标的同时扩展上下文会增加所需的集合大小,而其占上下文的分数在测试范围内下降。使用固定支持事实的实验表明,额外的背景信息会将其标记推低注意力排名并减少其注意力质量。对保留权重进行重新归一化可以显著减少所需集合大小,表明这也取决于所选表示的组合方式。条件理论模型解释了竞争和注意力质量保留如何在没有更多不同信息可检索的情况下产生不断增长的集合大小。这些结果为测量语言模型中的有效注意力集合大小并研究其对上下文、竞争和聚合的依赖性提供了一种方法。

英文摘要

How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑