时间递归有利于更少的层数
Temporal Recurrence Favors Fewer Layers
浏览论文内容
中文总结 AI 辅助
该研究将递归模型中的深度需求视为计算分配问题,发现在流式任务中,时间递归可将最优计算分配转向更少的层数,且性能相当或更好。
中文摘要 AI 辅助
在流式任务中,递归模型可以跨时间携带潜在计算,使得每次更新都能建立在先前产生的表示之上。这引发了一个基本问题:一旦时间递归提供了跨步骤的序列计算,每个步骤内部仍然需要多少深度?先前的工作表明,递归可以使浅层模型具有竞争力。我们则将这个问题作为一个计算分配问题来研究,在多个计算预算下,改变步骤内深度、专家宽度以及每层并行专家数量。对于每个预算,我们比较了最优的递归和非递归分配,以及在近似匹配的每步计算下它们所达到的性能。在Sokoban和自回归FineWeb语言建模中,我们发现时间递归将最优计算分配转向显著更少的层数,同时性能相当或更好。
英文摘要
In streaming tasks, recurrent models can carry latent computation across time, allowing each update to build on representations produced earlier. This raises a basic question: once temporal recurrence provides sequential computation across steps, how much depth is still needed within each step? Prior work has shown that recurrence can make shallow models competitive. We instead study this question as a compute-allocation problem, varying within-step depth, expert width, and the number of parallel experts per layer across several compute budgets. For each budget, we compare the best observed recurrent and non-recurrent allocations and the performance they achieve under approximately matched per-step computation. Across Sokoban and autoregressive FineWeb language modeling, we find that temporal recurrence shifts the best observed compute allocation toward substantially fewer layers, with comparable or better performance.
发表机构
- Mila – Québec AI Institute(米拉-魁北克人工智能研究所)
- Université de Montréal(蒙特利尔大学)
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。