发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过3x3数学推理实验,发现在线策略蒸馏中提示广度与回滚刷新存在交互作用,且提示效率受刷新和推理预算影响。
AI 中文摘要
在线策略蒸馏(OPD)需要多少提示,答案又如何依赖于生成其训练响应的学生策略?我们联合研究这两个控制因素:提示广度与回滚刷新。一个3x3的数学推理实验固定了14,080条轨迹和110次优化器更新,同时变化提示库和生成响应的策略快照数量。使用十个快照时,八个提示达到24.09%的平均准确率,接近14,080个不同提示的24.51%。然而,当响应冻结在初始策略时,增加广度将准确率从21.16%降至19.05%;在每次更新刷新下,增加广度将准确率从23.61%提升至25.57%。由此产生的交互效应为4.07个百分点,95%问题配对区间为[2.00, 6.28]。在两位教师下的匹配比较揭示了第二个反转:周期性模型在短预算准确率和答案完成度上更高,但冻结响应模型在32K输出限制下以1.7-1.8倍的响应令牌数在平均准确率上反超。这些结果表明,OPD中的提示效率可能同时依赖于刷新和推理预算。
英文摘要
In on-policy distillation (OPD), the student determines which states receive teacher supervision. As its policy evolves, earlier response prefixes become less likely even though teacher-student disagreement on them persists. Under matched trajectory and optimization budgets, neither more queries nor more frequent rollout resampling is uniformly beneficial. Current-policy rollouts outperform initial-policy rollouts at shorter response budgets, but this ranking reverses at longer budgets. Replaying initial-policy rollouts after current-policy training improves accuracy, whereas replaying fixed recent rollouts does not reproduce the gain. We propose R-OPD, a gradient-triggered curriculum that adaptively selects when to revisit initial-student trajectories. When changes in mean gradients fall within minibatch-level variation for two consecutive comparisons, training switches from the next iteration onward to initial-policy replay. Across eight mathematics benchmarks and three training-data orders, R-OPD improves average accuracy over continued current-policy sampling by 2.25/4.44 percentage points at 16K/32K for a 0.6B student and 2.81/6.15 points for a 1.7B student. With a 30B-A3B teacher, it improves 8B accuracy by 4.05 points at 32K. At 32K, R-OPD exceeds a fixed schedule of 40 current-policy updates followed by 20 replay updates by 1.39/1.66/1.85 points for 0.6B/1.7B/8B. At the same generation cap, R-OPD also produces longer responses, suggesting that well-timed revisits help students use more of their reasoning capacity.
Comments26 pages, including references and appendices