AI 中文总结
针对循环语言模型中LoRA共享更新的后循环偏差,提出Loop Dropout方法,通过随机掩蔽与逆生存重缩放保持期望更新强度,提升跨循环适配,增强数学推理、指令调优和代码生成性能。
AI 中文摘要
循环语言模型通过重复应用相同的Transformer块,将计算深度与参数数量分离。适配这些模型需要一种共享更新,该更新在隐藏状态随循环计算演变时保持有效。我们的实证分析揭示了标准低秩适配(LoRA)中存在显著的后循环偏差:共享更新在较后的循环位置更为有效。这种不平衡促使我们在不同的应用组合下训练共享更新。仅随机省略适配器应用并不能提高任务性能;它会降低训练期间预期的更新强度,同时使推理保持不变。我们引入了循环丢弃(Loop Dropout),它将适配器应用的随机掩蔽与逆生存重新缩放相结合,以保持预期的更新强度并促进跨循环的有效适配。大量实验表明,在模型规模、适配器秩和训练方案中,数学推理能力均得到提升,其益处扩展到通用指令调优和代码生成。循环丢弃优于现有的LoRA变体和适配器正则化器,进一步分析显示早期循环适配更强。每个骨干循环保持活跃,推理时在所有循环中应用适配器,使用标准LoRA,无需额外的可训练参数或推理计算。
英文摘要
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update provides limited adaptation at early loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation. Code is available at https://github.com/NUS-HPC-AI-Lab/loop-dropout .