arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

论门控的重要性:状态空间模型中的记忆化与上下文学习

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

William L. Tong, Aryo Lotfi, Emmanuel Abbe, Kostas Vaggelakos, Vishnu Banna, Etai Littwin, Josh Susskind, Cengiz Pehlevan, Eran Malach

arXiv 2609.16540首次发表:更新:

发表机构

Apple; Harvard University(苹果公司; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过理论和实验揭示,门控机制使状态空间模型偏向记忆化而延迟上下文学习,但有助于长序列泛化,为改进线性时间模型提供依据。

AI 中文摘要

状态空间模型(SSMs)已成为Transformer的一种引人注目的替代方案,能够以恒定内存和线性计算实现序列建模。尽管SSMs展现出合理的性能和有利的计算特性,但在需要上下文学习和精确检索的任务上,它们仍然落后于Transformer,这阻碍了它们在大规模语言建模中的采用。在这项工作中,我们证明,通过研究门控机制(现代循环网络中的一个普遍组件)的作用,可以解释SSMs在这些领域中的成功与失败。具体来说,我们通过理论和实验表明,这种门控机制导致SSMs首先学习一种权重内的“记忆化”解决方案,同时延迟甚至阻止其收敛到正确的上下文学习解决方案。重要的是,即使在架构或其内存容量没有根本限制的情况下,这种情况也会发生。另一方面,我们发现门控通常有助于改善对长序列长度的泛化能力。我们的结果阐明了门控机制在塑造SSMs的训练动态和泛化能力方面的关键作用,并为理解和改进线性时间模型提供了基础。

英文摘要

State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights "memorization" solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.

Comments25 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑