FactorEngram:面向语言模型的基于基级门控的分解N-gram记忆
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
浏览论文内容
中文总结 AI 辅助
FactorEngram通过基级门控的分解n-gram记忆,使相关模式共享组件并实现上下文调制,在340M和1B参数模型上提升语言建模和下游任务性能。
中文摘要 AI 辅助
基于查找的记忆已成为扩展大型语言模型(LLM)参数的一种有前景的方法。它检索局部标记模式(如n-gram)的学习表示,而不是通过连续的计算层重建它们。然而,现有设计如Engram将每个检索到的嵌入视为一个整体单元。每个嵌入存储在其自己的哈希槽中,并由单个标量门控调制。因此,多义模式无法选择性地读出与其上下文相关的记忆组件。此外,参数仅通过哈希冲突共享,这在很大程度上与语义无关。我们提出FactorEngram,一种具有基级上下文门控的分解n-gram记忆。FactorEngram在跨模式共享的基向量字典上检索稀疏正则化系数,因此相关模式可以重用公共组件。同一字典也用于门控。骨干隐藏状态与每个基向量进行评分,以在重建前门控相应的系数,从而让上下文单独调制每个记忆组件。FactorEngram还涵盖单个标记和多标记n-gram,并且我们系统地研究了记忆分支应插入的位置。在340M和1B参数的Transformer骨干上,FactorEngram改善了语言建模和下游任务性能。消融研究证实了每个组件的贡献,并确定在中间层的注意力子层之前插入是一种有效配置。
英文摘要
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
发表机构
- Baidu Inc.(百度公司)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。