arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15062cs.CLcs.LG

RecurrentGPT:通过Transformer中的循环调制实现表达性深度

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Amr Hegazy, Amr Alanwar, Mostafa Elhoushi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RecurrentGPT,通过循环调制解决Transformer语言模型表达性与内存效率的矛盾,在多约束下优于基线模型,实现参数与内存的高效利用。

中文摘要 AI 辅助

缩放Transformer语言模型时,表达性与内存效率之间存在固有矛盾:各层采用独特权重可保留功能专业化(从输入基础到抽象细化),但会产生大量内存占用;而标准深度共享会强制采用统一变换,导致表示多样性崩溃,降低建模质量。本文提出RecurrentGPT,这是一种循环深度Transformer,其中固定深度的前奏(prelude)和尾声(coda)块将单个共享核心包围,该核心会迭代R次。受门控循环神经网络启发,我们采用轻量级投影和逐元素更新门,该门以隐藏状态、固定前奏输出以及每步重新采样的噪声为条件,对循环更新进行调制。这使模型能在多次循环中对输入进行专业化处理,仅需少数几层即可实现功能多样性,而非需要大量独特层。在等FLOPS约束下,3层RecurrentGPT与12层GPT-2 Small基线的准确率相当,训练和推理FLOPs相近,且在所有9种按预算划分的规模单元中,在混合专家(MoE)和重尾深度采样方面表现领先;在中等和大规模下,它在标准token预算下接近密集模型质量,在中等规模下,当预算翻倍后则超过密集模型。在等参数(isoPARAMS)约束下,更深的循环在匹配参数和数据预算时,验证损失为2.76,而非循环对应模型的验证损失为2.84。我们的结果表明,自适应深度复用是一种用参数换取质量的合理策略:在大规模下,编译生成延迟增加10%,参数减少63%,峰值解码内存减少59%。

英文摘要

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.

发表机构

  • The German University in Cairo(开罗德国大学)
  • Technical University of Munich(慕尼黑工业大学)
  • Cerebras Systems Inc.(赛布拉斯系统公司)

机构由 AI 辅助整理,请以论文原文为准。

↑