发表机构
Beihang University; China Telecom eSurfing Cloud(北京航空航天大学; 中国电信天翼云)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出WDL-OPD混合约束协同训练方法,在Qwen3 17亿和40亿参数实验中,于四个规模-领域设置里获最强学生检查点,提升MATH500准确率,稳定代码生成表现,支持稳定性假设。
AI 中文摘要
在线策略蒸馏(OPD)在学生自身采样的轨迹上使学生与教师对齐,减少了离线蒸馏存在的训练-测试状态不匹配问题。不过,相同的反馈回路可能不稳定:每次更新都会改变策略以及下一次更新所计算的状态。我们提出WDL-OPD,这是一种带有两个可训练策略的混合约束协同训练方法。锚定策略生成所有轨迹,辅助策略评估相同的访问状态,通过反向KL散度将它们的token分布的几何混合与冻结教师匹配。两个策略都接收梯度。我们表明,冻结辅助策略可恢复与OPD²和W2S-OPD密切相关的锚定加对比代理目标,而联合训练会产生静态delta无法表达的分支级自由度。在记录的Qwen3实验中,规模为17亿和40亿参数时,WDL-OPD在四个规模-领域设置中均产生最强的学生检查点。它将MATH500的准确率从0.630提升至40亿参数时的0.685,从1.7亿参数时的0.521提升至0.585。在代码生成任务中,七个单策略OPD配置出现熵增长或轨迹退化,而协同训练达到独立重新评估的开发分数0.637和0.375。由于若干比较在课程或初始化方面存在差异,这些结果支持稳定性假设而非普遍因果主张。我们提供了测试该假设所需的精确训练算法、失败证据和受控比较矩阵。
英文摘要
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The same feedback loop can nevertheless be unstable: each update changes both the policy and the states on which the next update is computed. We introduce WDL-OPD, a mixture-constrained co-training method with two trainable policies. An anchor policy generates every rollout, an auxiliary policy evaluates the same visited states, and a geometric mixture of their token distributions is matched to a frozen teacher by reverse KL. Both policies receive gradient. We show that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express. In recorded Qwen3 experiments at 1.7B and 4B scale, WDL-OPD produces the strongest student checkpoint in each of four scale-domain settings. It raises MATH500 accuracy from 0.630 to 0.685 at 4B and from 0.521 to 0.585 at 1.7B. In code generation, seven single-policy OPD configurations exhibit entropy growth or trajectory degradation, while co-training reaches independently re-evaluated development scores of 0.637 and 0.375. Because several comparisons differ in curriculum or initialization, these results support a stabilization hypothesis rather than a universal causal claim. We provide the exact training algorithm, failure evidence, and the controlled comparison matrix needed to test that hypothesis.