发表机构
University of Illinois Urbana-Champaign; Princeton University; Westlake University(伊利诺伊大学厄巴纳-尚佩恩分校; 普林斯顿大学; 西湖大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析Qwen3-1.7B多教师同策略蒸馏中的梯度、优化器更新和精度影响,揭示了损失平均、Adam一阶矩、BF16舍入及KL梯度截断对参数更新和任务性能的作用。
AI 中文摘要
多教师同策略蒸馏(MOPD)旨在将经过强化学习(RL)训练的多个教师的能力整合到一个学生模型中,但教师信号如何影响参数变化仍未被充分探索。我们以Qwen3-1.7B为研究对象,使用四个领域教师,这些教师与学生模型从相同的初始化出发,通过RL训练得到,并比较了梯度、优化器更新和任务学习曲线,同时辅以SmolLM3-3B的诊断实验。我们发现多个因素影响教师信号。首先,损失平均隐式地对响应进行加权:词元平均偏向于更长的响应,而均衡各领域贡献则在该领域内保留了这种加权。其次,Adam的一阶矩减小了参数更新的差异:尽管原始梯度存在差异,教师之间的余弦相似度为0.83,而不同平均规则之间的余弦相似度为0.96。第三,BF16舍入隐藏了小的变化:约97%的FP32主权重与初始化不同,但只有7%至11%的BF16权重与初始化不同。最后,top-64交集KL梯度与Qwen的全词汇梯度高度匹配,但对任务性能的影响取决于平均方式:在响应平均下,数学准确率比采样词元策略梯度(PG)高2.6个百分点,而在全局词元平均下则低2.1个百分点。
英文摘要
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.