线性注意力如何记忆
How Linear Attention Remembers
AI总结:
本研究通过分析分解和因果干预,揭示了线性注意力中固定大小循环记忆如何通过集中式读写实现选择性事实召回,并指出其存在跨事实耦合、负载增加时性能下降及混合架构下依赖KV状态等特性。
AI中文摘要:
线性注意力用固定大小的循环状态取代了标准注意力中不断增长的键值(KV)缓存,从而在上下文长度增加时大幅减少内存增长。然而,这种效率改变了过去信息的存储方式:许多令牌必须共享并反复更新同一内存。我们研究了这种循环状态如何作为记忆系统运作。通过分析分解以及对预训练的GLA和GDN模型进行受控因果干预,我们追踪了被召回信息的写入、保留及后续访问过程。我们发现,特定事实的信息通过集中的、内容相关的写入进入循环记忆,并在之后通过集中的查询时读取路径被访问。多个事实可以在同一状态内保持选择性可访问,然而其内部表征表现出跨事实的因果耦合,而非类似KV的独立存储。随着记忆负载的增加,召回能力和定向可编辑性均会下降,而在所测试的范围内,仅经过的上下文长度的影响则要小得多。对后续写入的因果干预进一步表明,干扰是由其与现有记忆的重叠所塑造的。最后,在将循环层与完全注意力相结合的混合架构中,直接支持召回的实际运行内存主要转移到完全注意力的KV状态。总之,这些结果揭示了固定大小的循环记忆如何在共享存储的情况下支持选择性召回,同时暴露了将其与可寻址令牌的KV记忆区分开来的干扰和容量限制。
英文摘要:
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memory. We study how this recurrent state functions as a memory system. Using an analytical decomposition together with controlled causal interventions in pretrained GLA and GDN models, we trace how recalled information is written, retained, and later accessed. We find that fact-specific information enters recurrent memory through concentrated, content-dependent writes and is later accessed through concentrated query-time read pathways. Multiple facts can remain selectively accessible within the same state, yet their internal representations exhibit cross-fact causal coupling rather than independent KV-like storage. As memory load increases, both recall and targeted editability degrade, whereas elapsed context alone has a substantially smaller effect within the tested regime. Causal interventions on subsequent writes further show that interference is shaped by their overlap with existing memory. Finally, in hybrid architectures that combine recurrent layers with full attention, the runtime memory directly supporting recall shifts predominantly to the full-attention KV state. Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.