发表机构
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究策略蒸馏中局部监督受轨迹结果混淆的问题,引入结果解析诊断方法区分不同情况,通过数学推理研究等实验得出相关比例及结果,指出局部化限制,贡献在于诊断而非新训练方法。
AI 中文摘要
策略蒸馏(OPD)在学生自身轨迹上训练,教师在学生访问的前缀处提供密集的令牌级似然性。这些似然性通常在局部读取:一致似乎可以安全模仿,而不一致则似乎表明存在错误。我们表明这两种读取都受到完整轨迹结果的混淆。我们引入了一种结果解析诊断方法,将逐点师生差异与最终答案正确性交叉,区分安全模仿、有成效的差异、有害的差异和失败时的一致。在一项八种子的数学推理研究中,失败时的一致占汇总响应令牌质量的67.84%;使用Qwen2.5-7B/32B对时,这一比例仍为67.68%。该结果在阈值、序列级、格式和截断审核中都持续存在。即使在Qwen3教师在所有四次独立尝试中都能解决的提示上,学生准确率提高到86.91%,但失败时的一致仍为14.76%。然后我们进行了三个匹配的训练探针,使用可用信号模仿、掩盖或对比整个轨迹;没有一个能持续降低失败时的一致。结果指出了一个局部化限制:局部差异与轨迹级结果配对并不能确定失败的轨迹在何处变得无法恢复。解决这一限制需要额外的位置信息,如过程标签、教师从学生前缀的延续或跨展开的令牌级对齐。因此,我们的贡献是诊断性的,而不是一种新的训练方法。
英文摘要
On-policy distillation (OPD) trains a student on its own trajectories while a teacher supplies dense token-level likelihoods at student-visited prefixes. These likelihoods are often read locally: agreement appears safe to imitate, whereas disagreement appears to identify an error. We show that both readings are confounded by the outcome of the completed trajectory. We introduce an outcome-resolved diagnostic that crosses pointwise teacher-student divergence with final-answer correctness, separating safe imitation, productive divergence, harmful divergence, and agreement-on-failure. In an eight-seed mathematical-reasoning study with a Qwen3-8B student and Qwen3-32B teacher, agreement-on-failure constitutes 67.84% of pooled response-token mass; with a Qwen2.5-7B/32B pair it remains 67.68%. The result persists across threshold, sequence-level, format, and truncation audits. Even on prompts that the Qwen3 teacher solves in all four independent attempts, student accuracy rises to 86.91% but agreement-on-failure remains 14.76%. We then run three matched training probes that use the available signals to imitate, mask, or contrast whole trajectories; none consistently reduces agreement-on-failure. The result points to a localization limitation: local divergence paired with a trajectory-level outcome does not identify where a failed trajectory became unrecoverable. Addressing this limitation requires additional positional information, such as process labels, teacher continuations from student prefixes, or token-level alignment across rollouts. Our contribution is therefore diagnostic rather than a new training method.