从修复推理中学习:根因引导的在线策略蒸馏
Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对在线策略自蒸馏中参考引导与所需修正不匹配及蒸馏陷阱问题,提出根因引导的在线策略蒸馏(RC-OPD),通过定位学生推理最早错误并迭代修复,实现根因与锚点双重引导,显著提升性能。
中文摘要 AI 辅助
在线策略自蒸馏(OPSD)使用参考解作为特权后见之明,来监督学生生成的推理轨迹。然而,基于参考的引导可能解释了一个正确的解决方案,却没有解决学生自身推理为何失败的问题。所提供的引导与所需修正之间的这种推理不匹配,可能鼓励学生借用正确的结论,而使其推理错误得不到解决。此外,在整个轨迹中应用相同的后见之明可能会陷入蒸馏陷阱,即对有效推理的不必要约束与对实质性错误的修正相互竞争。为解决这些问题,我们提出了根因引导的在线策略蒸馏(RC-OPD),该方法利用学生自身推理的修复来提供针对其特定错误的引导,同时建立在有效进展之上。对于每次失败的尝试,RC-OPD定位最早的实质性错误,制定局部修正,并将修正后的中间结果作为有效前缀的锚点。一个迭代的诊断-修复-继续过程通过学生的继续来测试修复,在固定的修复预算内识别进一步的错误。对于达到正确答案的修复链,根因引导蒸馏使用失败诊断和修正目标来监督错误片段,而锚点引导蒸馏则用通向修复后中间结果的推理链来支持相应的有效前缀。我们在多个数据集和模型规模上评估RC-OPD。大量实验和分析表明,它缓解了推理不匹配和蒸馏陷阱,带来了显著的性能提升。
英文摘要
On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student's own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis--repair--continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root--cause--guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.
发表机构
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
- School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
- School of Information Technology and Management, University of International Business and Economics(对外经济贸易大学信息与技术管理学院)
机构由 AI 辅助整理,请以论文原文为准。