发表机构
Meta AI; University of California, Riverside(Meta AI; 加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出递归框架,通过动态共同进化(DCE)和自精炼简洁学习(SRCL)改进同策略自蒸馏,使教师与学生共同进化,在数学基准上显著超越OPSD。
AI 中文摘要
同策略蒸馏(OPD)通过让学生模型生成轨迹,然后将其下一个词元预测与外部教师模型的预测进行匹配来训练学生模型。这为学生提供了密集的词元级监督。同策略自蒸馏(OPSD)消除了对外部教师的需求。具体而言,学生模型的第二个冻结副本(在其上下文中给定真实答案)充当教师。学生模型仅接收问题,并学习模仿特权教师模型,而教师在整个训练过程中保持冻结。先前的工作表明,冻结教师有助于训练稳定性,但我们认为这可能会阻止教师纳入学生在训练期间学到的改进。我们的主要贡献是通过一个基于两个互补组件的递归框架来解决这一局限性。首先,我们让特权教师与学生共同进化,使得在一轮中学习的修正可以指导下一轮,我们将这一过程称为动态共同进化(DCE)。其次,由于更强的修正也可能使回答过于冗长和自我批评,我们还额外训练模型自身同策略回答的较短且经过验证的改写版本。我们将这一互补目标称为自我精炼简洁学习(SRCL)。总体而言,我们的综合评估表明,DCE+SRCL在多个模型规模和四个竞赛级数学基准上均优于OPSD。具体而言,在Qwen3-8B上,DCE+SRCL达到65.97%的Average@12,比OPSD高出35.62个百分点,同时相对于单独使用DCE,平均输出长度减少了7.80%。
英文摘要
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.