arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从演化错误中学习:用于在线策略蒸馏的自适应迭代修复

Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

Rui Li, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu

arXiv 2610.02700首次发表:更新:

发表机构

University of Science and Technology of China; Nanyang Technological University(中国科学技术大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自适应迭代修复框架AIR-OPD,通过错误到修复的监督改进在线策略蒸馏,在数学推理基准上显著提升性能。

AI 中文摘要

在线策略自蒸馏(OPSD)在从学生自身策略采样的轨迹上提供密集的令牌级反馈,这是一种比强化学习的结果级奖励更丰富的训练信号。该反馈来自一个以完整参考解决方案为条件的教师模型,而该参考解决方案是学生无法获得的。参考解决方案指定了目标,但并未说明如何从学生当前的错误走向目标,从而产生了解决方案条件的捷径风险。我们引入了AIR-OPD,一种用于在线策略蒸馏的自适应迭代修复框架,提供从错误到修复的监督。给定一个失败的响应,一个引导生成器为当前错误合成修复引导。学生利用该引导进行在线策略重试。如果重试仍然不正确,生成器会为新观察到的错误生成新的修复引导。在每一轮中,一个固定的教师模型将引导作为特权上下文接收,并在学生最新失败响应的错误对齐区域上监督学生。结果感知的阶段加权有利于早期修复阶段,并奖励其即时重试通过验证的阶段。我们在DAPO-Math-17K数据集上训练AIR-OPD,并在AIME24、AIME25和HMMT25上评估,同时在MMLU-Pro和GPQA上进行分布外测试。我们考察了两种引导来源:来自当前学生策略的自我引导和来自更大模型的外部引导。对于Qwen3-4B和Qwen3-8B,AIR-OPD均取得了最佳的数学推理平均分,比最强基线提高了最多3.6分,同时在分布外基准上保持了基础模型的性能。

英文摘要

On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.

Comments21 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑