arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18230cs.LG

在循环语言模型中分配循环计算

Allocating Recurrent Compute in Looped Language Models

Ruhai Lin, Yiyang Guo, Rui-Jie Zhu, Hao Ye, Jason K. Eshraghian

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对循环语言模型,提出仅重复混合器、单次执行密集FFN的MixerLoop,在不同参数规模下验证其性能,可保留循环深度优势并降低计算量。

中文摘要 AI 辅助

循环语言模型通过重复应用共享计算来提升推理和知识操作能力。现有系统通常会重复整个层堆栈,尽管混合器(mixer)和密集前馈网络(FFN)执行不同操作且成本各异。我们提出一个更具体的问题:应当循环什么?我们将循环视为状态更新的重复组合,并认为当应用程序在任务读出时能暴露出新的跨位置影响方向且该方向仍可观测时,其具有价值。迭代传输秩(ITR)描述累积影响轨迹;边际ITR描述连续应用所贡献的非冗余影响。这一观点催生了MixerLoop,即重复每个门控DeltaNet混合器,仅执行其密集FFN一次。我们在15M和110M参数规模下,于相同数据、初始化和架构条件下,将MixerLoop与无循环、全块循环进行对比。有限上下文干预测试后续混合器应用是否在最终语言模型读出时产生显著、非冗余且有益的变化。MixerLoop在15M参数下的聚合CORE指标优于FullLoop,在110M参数下保留了FullLoop 41.5%的CORE提升,同时将循环骨干投影FLOPs减少了45.9%。这些结果表明,无需重复执行密集FFN,即可保留循环深度的优势。

英文摘要

Looped language models improve reasoning and knowledge manipulation by applying shared computation repeatedly. Existing systems usually repeat an entire layer stack, although a mixer and a dense feed-forward network (FFN) perform different operations and have different costs. We ask a narrower question: what should loop? We view recurrence as repeated composition of a state update and argue that an application is valuable when it exposes a new cross-position influence direction that remains observable at the task readout. Iterative Transport Rank (ITR) describes the cumulative influence trajectory; marginal ITR describes the nonredundant influence contributed by successive applications. This view motivates MixerLoop, which repeats each Gated DeltaNet mixer while applying its dense FFN once. We compare MixerLoop with no recurrence and full-block recurrence at 15M and 110M parameters under the same data, initialization, and architecture. A finite context-off intervention tests whether later mixer applications produce distinct, non-negligible, and beneficial changes at the final language-model readout. MixerLoop surpasses FullLoop on aggregate CORE at 15M and retains 41.5% of its CORE improvement at 110M while reducing recurrent-backbone projection FLOPs by 45.9%. These results show that the benefits of recurrent depth can be retained without repeatedly executing the dense FFN.

发表机构

  • University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

↑