arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

好老师因材施教:联合在线学习与教学

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Randy Ardywibowo, Arnav Dalal, Jiantao Jiao

arXiv 2610.10447首次发表:更新:

发表机构

Perplexity; NVIDIA(Perplexity; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出JOLT方法,通过联合训练特权教师与非特权学生,利用KL正则化目标与在线策略蒸馏,解决强化学习稀疏奖励问题,提升训练效率与性能。

AI 中文摘要

从结果奖励中进行强化学习(RL)会遭受稀疏监督的困扰,尤其是在困难、长视界任务中,成功的轨迹稀少且生成成本高昂。在线策略蒸馏(OPD)提供了一种有吸引力的替代方案,它通过沿着学生自身生成轨迹,从更强的教师那里提供密集的令牌级监督。自蒸馏方法进一步消除了对单独教师模型的需求,通过让同一策略基于特权信息进行条件化,使其充当自己的教师。然而,仅凭特权条件化并不能保证所得的蒸馏更新能改善学生。实际上,特权信息可能导致教师通过学生无法获得的捷径来解决问题,从而产生与学生当前行为不匹配的监督。因此,即使是表现更好的教师也可能提供导致学生性能下降的指导。为了解决这个问题,我们分析了特权教师的选择如何影响学生的更新。我们推导出教师局部蒸馏更新成为学生奖励梯度的正倍数的充要条件。我们的分析表明,教师不仅应在任务上表现良好,还应提供适合学生当前能力的指导。这一特征激发了一个实用的教师训练代理目标,该目标将结果奖励与对学生进行令牌级KL散度正则化相结合。基于这一结果,我们提出了联合在线学习与教学(JOLT),该方法以两种角色联合训练单一策略:使用KL正则化目标训练的特权教师,以及使用密集在线策略蒸馏训练的非特权学生。在数学推理、编程、工具使用和终端使用方面,JOLT提高了训练效率和性能,并且通过学生奖励进一步获得了增益。

英文摘要

Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑