发表机构
The Hong Kong University of Science and Technology (Guangzhou); University of Science and Technology of China; Alibaba Group; Stanford University(香港科技大学(广州); 中国科学技术大学; 阿里巴巴集团; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型推理修正,提出分段式策略内蒸馏(Seg-OPD),通过不确定性选择学生分段并用教师改写进行监督,提升推理准确率,平均相对提升5.22%。
AI 中文摘要
策略内蒸馏(OPD)通过利用教师模型提供的密集逐词监督,在学生的自生成轨迹上训练,从而提升大语言模型的推理能力。然而,逐词的OPD并未明确提供一种连贯的替代推理步骤,以展示如何修正学生的步骤来改善后续推理。此外,当学生产生退化的推理前缀时,这种范式可能变得效果不佳,因为后续的教师监督仍以该前缀为条件,可能强化不良的推理模式。在本工作中,我们专注于通过分段式策略内蒸馏(Segment-wise OPD)学习推理修正,以重写中间推理步骤并更好地支持后续推理。通过受控的推理干预实验,我们发现用教师的改写替换学生分段能提升后续推理的准确性。因此,我们解决了将教师改写转化为明确监督以学习推理修正的问题。我们提出了分段式策略内蒸馏(Seg-OPD),该方法基于不确定性指标选择学生分段,并获取相应的教师改写。Seg-OPD训练学生偏好教师改写而非配对的学生分段,同时保留密集的逐词OPD监督。在数学推理和竞争性编程任务上的大量实验表明,经过Seg-OPD训练的学生在修正成功率上优于基线方法。Seg-OPD在推理准确性上持续优于所比较的最先进基线,在多种模型和任务上平均相对提升5.22%。代码可在该https URL获取。
英文摘要
On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revised to improve subsequent reasoning. Furthermore, this paradigm can become less effective when the student produces a degenerate reasoning prefix, as subsequent teacher supervision remains conditioned on that prefix and may reinforce poor reasoning patterns. In this work, we focus on learning reasoning revision with segment-wise OPD to rework intermediate reasoning steps and better support subsequent reasoning. Through controlled reasoning interventions, we find that replacing student segments with teacher redrafts improves subsequent reasoning accuracy. Therefore, we address the problem of turning teacher redrafts into explicit supervision for learning to revise reasoning. We propose Segment-wise On-Policy Distillation (Seg-OPD), which selects student segments based on an uncertainty metric and obtains corresponding teacher redrafts. Seg-OPD trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise OPD supervision. Extensive experiments on mathematical reasoning and competitive programming tasks show that Seg-OPD-trained students achieve higher revision success rates than baselines. Seg-OPD consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks. Code is available at https://anonymous.4open.science/r/Seg-OPD.