arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17251cs.CL

Transformer层间的持久循环记忆——提升语言模型泛化能力

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

  • FITec Labs / Ericsson São Paulo(FITec实验室 / 爱立信圣保罗)

机构由 AI 辅助整理,请以论文原文为准。

Eduardo Novaes Hering

AI总结:

本文提出在Transformer层间插入带GRU的持久循环记忆模块,仅增3.7%参数即显著降低28.5%的评估损失,减少过拟合,证明该拓扑是提升小模型泛化的简单有效方法。

AI中文摘要:

我们提出了一种对仅解码器Transformer的简单架构修改:一个持久的循环状态,通过交叉注意力观察隐藏表示,通过GRU更新自身,并通过门控加法调节后续处理。将该模块插入6层Transformer的下半部分和上半部分之间,仅增加3.7%的参数,同时将评估损失从2.438±0.004降至1.743±0.018,对应于在保留的语言建模数据上28.5%的降低。该改进在5个随机种子下具有统计显著性(p<0.01),并对应于过拟合的减少(泛化差距0.12对0.26)。通过受控消融实验,我们证明该改进完全源于持久记忆拓扑,而非辅助自预测目标。具有相同拓扑但无辅助损失的模型表现相当,而随机辅助损失则无任何益处。表示探测显示,持久状态编码了叙事位置(52%对33%的随机水平)——这是标准注意力难以高效维护的信息。我们的结果表明,用轻量级循环记忆桥接Transformer层是一种在小规模语言模型中提升泛化能力的简单有效的方法。

英文摘要:

We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing via gated addition. Inserted between the lower and upper halves of a 6-layer transformer, this module adds only 3.7\% additional parameters while reducing evaluation loss from $2.438 \pm 0.004$ to $1.743 \pm 0.018$, corresponding to a 28.5\% reduction on held-out language modeling data. The improvement is statistically significant across 5 random seeds ($p < 0.01$) and corresponds to reduced overfitting (generalization gap 0.12 vs 0.26). Through controlled ablations, we demonstrate that the improvement stems entirely from the persistent memory topology, not from auxiliary self-prediction objectives. A model with identical topology but no auxiliary loss performs equivalently, while a random auxiliary loss provides no benefit. Representation probing reveals that the persistent state encodes narrative position (52\% vs 33\% chance level)---information that standard attention maintains less efficiently. Our results suggest that bridging transformer layers with a lightweight recurrent memory is a simple, effective approach to improving generalization in small-scale language models.

↑