arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Proteus:面向长上下文序列建模的增量式内存激活机制

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville

arXiv 2608.16844首次发表:更新:

发表机构

Mila; Google(米拉计算科学研究所; 谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Proteus 是一种增量式内存激活机制,可嵌入多种神经内存架构,应用于 SWLA 等模型后,能在长上下文相关任务上实现性能提升,证明调度有效容量对序列建模的适用性。

AI 中文摘要

基于注意力机制的序列模型在处理长上下文时会产生二次计算成本,这推动了越来越多研究转向可将上下文压缩为紧凑状态的内存型模型。然而,现有大多数内存模型在整个序列处理过程中都采用静态内存,由于早期 token 无需承受压缩压力,会占用过多自由度,进而“污染”内存状态,导致后续上下文可用容量不足,还会增加已存储内容与后续到达内容之间的干扰。本文研究了一种新范式——增量式内存激活,即内存的有效容量会随上下文长度增长而逐步扩展。设置早期瓶颈可迫使模型更有效地压缩历史信息,而随时间解锁新容量则能减少干扰,提升对后续上下文的保留能力。我们将该范式实例化为 Proteus,这是一种可嵌入到多种神经内存架构中的简单机制,无需额外成本。我们将 Proteus 应用于 SWLA、Comba、Titans、Hope-Attention 等当前最优模型,在标准语言建模、推理任务,以及长上下文检索和理解任务上均观察到一致的性能提升,且随着上下文长度增加,增益也会增大。总体而言,我们的结果表明静态内存并非最优选择,而调度有效容量是序列建模中一种简单且广泛适用的工具。

英文摘要

The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throughout the entire sequence. Because early tokens face no compression pressure, they occupy too many degrees of freedom and "pollute" the memory state, leaving little capacity for later context and increasing interference between what is stored and what arrives next. We study a new paradigm of incremental memory activation, where the effective capacity of memory is progressively expanded as the context grows. Imposing an early bottleneck forces the model to compress history more effectively, while unlocking fresh capacity over time reduces interference and improves retention of later context. We instantiate this paradigm in Proteus, a straightforward mechanism that can be incorporated into a broad class of neural memory architectures at no additional cost. We apply Proteus to state-of-the-art models, including SWLA, Comba, Titans, and Hope-Attention, and observe consistent improvements on standard language modeling and reasoning, as well as on long-context retrieval and understanding, with gains that grow at longer context lengths. Overall, our results show that static memory is suboptimal and that scheduling effective capacity is a simple and broadly applicable tool for sequence modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑