发表机构
MBZUAI; The Chinese University of Hong Kong(穆罕默德·本·扎耶德人工智能大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过受控实验系统研究了循环语言模型中循环的有效性条件,发现循环可提升超训练视野的推理但损害知识,并提出通道级历史状态注入结合时间步条件作为更优设计。
AI 中文摘要
循环语言模型(LoopLMs)通过参数共享增加计算深度,为在不增加参数的情况下扩展推理计算提供了一条途径。然而,目前尚不清楚额外循环何时有益,以及架构选择如何影响其有效性。通过受控实验,我们系统地考察了(1)循环何时有帮助,(2)应将其应用于何处,以及(3)其条件设置如何影响性能。我们的评估涵盖了在知识和推理任务下,推理预算低于、等于和超过训练视野的情况。(1)我们发现,循环可以在训练视野之外改善推理性能,但同时会降低知识任务的表现,而更难的推理实例并不总是受益更多。(2)性能还取决于不同层和循环迭代的分配方式,表明仅凭有效深度不足以预测行为。非循环输出层提高了对欠展开的鲁棒性,而输入和输出层的首选位置随推理预算而变化。(3)最后,我们发现传统的初始状态注入对变化的循环深度提供的鲁棒性有限。因此,我们提出历史状态注入作为替代方案,并表明结合时间步条件设置的通道级历史状态注入提供了一种低成本且更有效的设计,在扩展展开时更好地保持知识,同时提高跨推理预算的鲁棒性。总体而言,我们的结果阐明了循环计算何时有帮助、何处失效,并为在不同推理预算下设计LoopLMs提供了实用指南。
英文摘要
Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.
CommentsPreprint, under-review