再循环
Recirculation
浏览论文内容
中文总结 AI 辅助
该研究提出推理时架构增强技术“再循环”及其自适应变体,无需额外训练,可显著降低Gemma3系列模型的困惑度、提升推理任务准确率,为架构演进提供新方向。
中文摘要 AI 辅助
我们描述了一种针对现成基础模型的推理时架构增强方法,该方法可显著降低困惑度并提升生成与推理任务的准确率。我们的方法在生成阶段几乎不会增加延迟,不过在预填充阶段需要串行处理。受前馈Transformer的状态更新受模型深度限制这一基本局限的启发,我们的技术“再循环(recirculation)”引入了一种特定形式的循环,使模型能作为动力系统运行并跟踪信念状态。我们将该技术与思维链计算(后者更适合复杂推理而非基础状态跟踪)、流行的深度循环技术(循环)以及代价高昂的循环Transformer训练区分开来。我们还提出并评估了一种自适应再循环变体,该变体仅需对超参数进行轻量调整,同时冻结原始模型权重。与现成基线相比,自适应再循环在Gemma3系列上取得了显著提升,包括在一组数据集上困惑度降低23%、在GSM8k上准确率提升21%,并在其他下游任务上实现了可靠的准确率提升。我们这种无需训练的方法通过利用模型自身特性来指导架构修改取得成功,为架构演进提供了一条路径——该路径由训练后网络的属性引导,而非受限于人为的任意设计选择。
英文摘要
We describe an inference-time architectural enhancement for off-the-shelf foundation models that systematically reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency during generation, though it requires serial processing in the prefill phase. Motivated by the fundamental limitation that state updates in feedforward transformers are bounded by model depth, our technique, recirculation, introduces a specific form of recurrence that allows the model to act as a dynamical system and track belief states. We distinguish this technique from chain-of-thought computation---which is better reserved for complex inferences rather than basic state tracking---as well as from popular depth-recurrence techniques (looping) and the costly training of recurrent transformers. We also propose and evaluate an adaptive variant of recirculation which requires only light tuning of hyperparameters while freezing the original model weights. Relative to the off-the-shelf baseline, adaptive recirculation achieves surprising gains on the Gemma3 family, including a consistent reduction in perplexity on a suite of datasets, a 21% increase in relative accuracy on GSM8k, and reliable improvements in accuracy on other downstream tasks. Our training-free approach succeeds by leveraging the model itself to inform architectural modifications, suggesting a route to architectural evolution guided by a trained network's properties rather than forced, arbitrary design choices.
发表机构
- Google DeepMind(谷歌DeepMind)
- University of Texas, Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。