AI 中文总结
ReOrder-OPD 提出以代理指标估计提示级教师延续可靠性并排序提示,结合 vanilla OPD 训练,在 Qwen3、Gemma4 等多模型数学及代码任务中提升了在线策略蒸馏效果。
AI 中文摘要
在线策略蒸馏(On-Policy Distillation, OPD)对学生智能体生成的轨迹应用 token 级别的教师监督,但该监督并非始终可靠。现有方法采用局部置信度或师生一致性对采样轨迹进行加权、过滤或截断,这些信号无法直接判断教师能否基于学生前缀生成正确答案,且轨迹级干预会将某一轮次的不可靠性与其提示的低预期训练价值混淆。我们将提示级教师延续可靠性 $R$ 定义为教师从学生前缀生成正确答案的概率,该概率在当前学生生成的所有前缀和轨迹上取平均。Oracle 实验表明,高 $R$ 提示能带来更大的 OPD 增益,且在固定提示池中,按 $R$ 降序训练的效果优于随机和升序排序。由于估计 $R$ 需要大量教师延续,我们采用独立学生轮次与验证器校正的同提示教师轨迹之间的最大 ROUGE-5 F1 作为代理指标。在该实际得分的 10 个等频区间中,平均 $R$ 单调上升,表明该代理可区分粗略的可靠性水平。ReOrder-OPD 按该代理对提示排序,再为 vanilla OPD 抽取独立的在线策略训练轨迹。它在 Qwen3 数学设置、Gemma4 数学设置及 Qwen3 代码设置中,所有匹配的聚合比较均有提升;在全部 6 个 FiRe-OPD 和 ExOPD 设置中的增益表明,提示排序与轨迹内监督具有互补性。
英文摘要
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its prompt. We define prompt-level teacher continuation reliability $R$ as the teacher's probability of reaching a correct answer from a student prefix, averaged over prefixes and trajectories induced by the current student. Oracle experiments show that high-$R$ prompts yield larger OPD gains and that descending-$R$ training outperforms random and ascending orders on a fixed prompt pool. Because estimating $R$ requires many teacher continuations, we use the maximum ROUGE-5 F1 between one independent student rollout and verifier-correct same-prompt teacher trajectories. Across ten equal-frequency bins of this actual score, mean $R$ rises monotonically, showing that the proxy separates coarse reliability levels. ReOrder-OPD sorts prompts by the proxy, then draws independent on-policy training trajectories for vanilla OPD. It improves every matched aggregate comparison across Qwen3 and Gemma4 mathematics settings and Qwen3 code settings. Gains in all six FiRe-OPD and ExOPD settings show that prompt ordering complements within-trajectory supervision.