arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环自改进:面向循环语言模型的动态跨循环在线蒸馏

Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

Yi Wang, Rui Qian, Yu Li, Haoyang Yao, Wenjie Wang

arXiv 2610.10623首次发表:更新:

发表机构

ShanghaiTech University; Fudan University; Southeast University; Peking University(上海科技大学; 复旦大学; 东南大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对循环语言模型后训练的局限,提出LoopOPD及动态版本D-LoopOPD,利用自身循环计算作监督,在Ouro-Thinking模型上提升了数学推理等能力。

AI 中文摘要

循环语言模型(LoopLMs)通过在循环计算步骤中复用共享参数,提供了一种参数高效的推理扩展方法。尽管其前景可观,但LoopLMs的有效后训练仍具挑战性。现有方法要么提供基于奖励的监督,这类监督在跨循环时稀疏且扩展成本高,要么依赖外部教师或特权信息,导致教师可用性有限或师生上下文不匹配。为解决这些局限,我们提出LoopOPD,一种跨循环在线蒸馏框架,它利用LoopLM内部额外的循环计算作为自身的监督源。LoopOPD使用冻结的终端循环策略作为计算特权教师,作用于中间循环学生生成的rollout,在无需外部教师或特权信息的情况下提供密集监督。我们进一步提出动态LoopOPD(D-LoopOPD),它会随着共享模型参数的更新持续刷新终端循环教师,实现循环自改进。我们刻画了蒸馏更新如何跨循环深度传播,并推导了单次更新能同时在两个循环深度产生局部改进的充分条件。在Ouro-Thinking模型上的实验表明,LoopOPD提升了数学推理能力,而D-LoopOPD通过动态教师更新进一步带来增益。尽管仅在数学数据上训练,所得模型在通用推理和代码生成基准上也有所提升,证明循环计算可作为LoopLMs的有效监督源。我们的代码和模型检查点将在论文接收后发布。

英文摘要

Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly to extend across loops, or rely on external teachers or privileged information, leading to limited teacher availability or teacher-student context mismatch. To address these limitations, we introduce LoopOPD, a cross-loop on-policy distillation framework that uses additional recurrent computation within a LoopLM as its own source of supervision. LoopOPD uses a frozen terminal loop policy as a compute privileged teacher for an intermediate loop student on student generated rollouts, providing dense supervision without an external teacher or privileged information. We further propose Dynamic LoopOPD (D-LoopOPD), which continually refreshes the terminal loop teacher as the shared model parameters are updated, enabling recurrent self-improvement. We characterize how distillation updates propagate across loop depths and derive sufficient conditions under which a single update yields simultaneous local improvement at both loop depths. Experiments on Ouro-Thinking models show that LoopOPD improves mathematical reasoning, while D-LoopOPD yields further gains through dynamic teacher updates. Despite being trained only on mathematical data, the resulting models also improve on general reasoning and code generation benchmarks, demonstrating that recurrent computation can serve as an effective source of supervision for LoopLMs. Our code and model checkpoints will be released upon acceptance.

Comments28 pages, 8 figures. Submitted to ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑