发表机构
University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统比较在线与离线策略蒸馏,发现词元级KL方向比回滚策略更影响性能,前向KL稳健而反向KL敏感,在线数据仅在特定条件下提升泛化。
AI 中文摘要
在线策略学习被认为可以减少灾难性遗忘、产生更稀疏的参数更新并改善泛化能力。然而,现有的监督微调与强化学习之间的比较同时变化了多种因素,使得回滚策略的贡献难以单独分离。我们在受控的强到弱蒸馏设置中研究回滚策略的影响,通过独立变化回滚策略、词元级KL方向和学习率,跨越Llama3和Qwen2.5模型家族以及涵盖科学、医学和算术领域的推理任务。我们的分析揭示了蒸馏动态的细致图景,其中回滚策略不一定扮演核心角色。相反,词元级KL方向更清晰地塑造了任务性能和输出覆盖范围,而学习率则控制遗忘和更新稀疏性。对KL梯度的分析和沿连续的学生-教师回滚策略谱的实验解释了这一模式:前向KL对回滚策略极为稳健,其性能稳定且强大,尽管回滚策略发生变化;而反向KL则更为敏感,并偏好学生生成的回滚。然而,在线策略数据在两种KL方向下都改善了对Countdown算术任务更难变体的泛化,尽管这一优势在后续的RLVR之后并不可靠地持续存在。我们的更广泛结论在去除梯度裁剪、使用采样KL估计器以及在需要更长推理链的任务上训练时依然稳健。总体而言,我们的结果挑战了在线策略回滚固有更优的观点,并表明其价值关键取决于目标、评估设置和优化超参数。
英文摘要
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.