发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReTeach 是仅用自身生成尝试与结果级验证的反思式自蒸馏框架,经多轮反思重试构建自教师,在六个基准上较 GRPO 提升平均准确率 1.39 个百分点。
AI 中文摘要
自蒸馏可在无需单独训练的更强大教师的情况下提升推理能力,但其有效性取决于自教师如何获得优于学生的优势。以参考答案或解决方案为条件可提供此类优势,但该信息可能不可用。反思提供了一种从自身生成的尝试中推导明确错误诊断与修正指导的方式,然而现有基于反思的方法常将其与参考信息、丰富任务反馈或持久记忆结合。我们提出ReTeach,一种反思式自蒸馏框架,仅利用自身生成的尝试与结果级验证,通过多轮反思与重试构建其自教师。从一次不成功的学生 rollout(展开)开始,教师在明确反思与重新尝试之间交替,直至成功或重试预算耗尽,无需参考答案或解决方案、外部诊断反馈或跨示例记忆。每次失败的重试为后续反思提供信息,而成功的修正为所得教师上下文的潜在效用提供结果级证据。一种感知结果的选择与加权策略区分初始正确、经反思修正及未解决的示例,为其类别归一化的蒸馏损失分配单独权重。通过在策略蒸馏,学生在自身展开的前缀处匹配教师的上下文条件 token(令牌)级预测分布,传递迭代修正的益处同时保留单步推理。在涵盖数学推理、科学问答与工具使用的六个基准上,ReTeach 较 GRPO 提升了平均准确率 1.39 个百分点。
英文摘要
Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student. Conditioning the teacher on reference answers or solutions can provide such an advantage, but this information may be unavailable. Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference information, rich task feedback, or persistent memory. We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification. Starting from an unsuccessful student rollout, the teacher alternates explicit reflection with renewed attempts until success or the retry budget is exhausted, without reference answers or solutions, external diagnostic feedback, or cross-example memory. Each failed retry informs subsequent reflection, while successful correction provides outcome-level evidence for the potential utility of the resulting teacher context. An outcome-aware selection and weighting strategy distinguishes initially correct, reflection-corrected, and unresolved examples, assigning separate weights to their category-normalized distillation losses. Through on-policy distillation, the student matches the teacher's context-conditioned token-level predictive distributions at prefixes of its own rollouts, transferring the benefits of iterative correction while retaining single-pass inference. Across six benchmarks spanning mathematical reasoning, science question answering, and tool use, ReTeach improves average accuracy over GRPO by 1.39 percentage points.