arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39661cs.CL

大型语言模型中注意力的演变:机制、权衡与新兴趋势

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Zhentao Tan, Jingyi Shen, Yanbo Li, Yao Liu, Yue Wu, Jieping Ye

首次发表
浏览论文内容

中文总结 AI 辅助

本综述通过五维视角分析大型语言模型中注意力机制的演变,揭示显式记忆与循环方法融合、跨层协调及多维记忆路由假设,强调高效架构设计正转向上下文记忆的组织与选择性使用。

中文摘要 AI 辅助

自注意力机制赋予大型语言模型对上下文进行细粒度、查询相关的访问能力,但密集的标记交互导致二次方的预填充成本以及随上下文长度增长的键值缓存。因此,研究涵盖了显式记忆压缩、稀疏访问、循环状态构建、结构化状态动力学以及异构机制组合。本综述将这些发展视为模型内部的上下文记忆进行分析。我们引入了一个五维视角——记忆表示、记忆更新、访问、读取和整合——描述了什么被表示、如何变化、什么有资格被查询、如何被读取以及读取结果如何形成输出。该视角在不强加单一计算模型的情况下比较了重叠的研究方向。我们利用来自14个主要模型系列和11个高性能开放权重端点的59条发布级记录,重构了机制层面的发展和架构采用情况。首先,显式记忆和循环状态方法保持不同的接口,但日益控制重叠的记忆功能。其次,异构架构日益在网络深度上进行协调:逐层组合将互补的记忆处理分布在表示阶段,而跨层重用将选定的记忆和路由伪影向前传递。因此,深度成为构建和管理上下文记忆的一个维度。第三,这些发展催生了一个有状态的多维记忆路由假设:持久记忆按时间范围、网络深度、基质类型和表示粒度进行组织,而协调的稀疏写入和稀疏读取决定了维护什么以及每个查询贡献什么。总体而言,高效的序列架构设计日益关注上下文记忆的组织、生命周期和选择性使用,而非孤立的注意力算子。

英文摘要

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.

发表机构

  • Alibaba Token Hub, Alibaba Group(阿里巴巴集团,阿里巴巴Token Hub)

机构由 AI 辅助整理,请以论文原文为准。

↑