发表机构
Beijing University of Posts and Telecommunications; Kuaishou Technology(北京邮电大学; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出全局状态模型(GSM),一种因果编码器-解码器架构,通过编码阶段集中选择与聚合长距离信息形成固定窗口共享状态,使解码器每步注意力成本与KV缓存不随历史长度增长,在保持性能的同时提升计算效率并降低缓存开销。
AI 中文摘要
高效的语言模型不仅要降低对过去上下文进行单次访问的成本,还要减少在各层中反复选择和处理历史信息的开销。我们引入了全局状态模型(GSM),这是一种因果编码器-解码器架构,将长距离信息的选择和聚合集中在编码阶段。通过多阶段的历史检索,编码器逐步将长距离信息整合到近期位置的表示中,形成一个具有固定窗口大小的共享状态。每个解码器层使用由前一层更新的查询来访问这一相同状态,从而在保持计算深度的同时,避免了重复构建历史键值(KV)表示和长距离索引。因此,解码器每步的注意力成本及其KV缓存大小均不随历史长度增长。实验表明,GSM在保持模型性能和使用长距离信息能力的同时,提高了计算效率并减少了缓存开销,为高效语言建模提供了一种共享状态架构。
英文摘要
Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.