发表机构
University of Illinois; UC San Diego; Rochester Institute of Technology; Clemson University; Miami University(伊利诺伊大学; 加州大学圣迭戈分校; 罗切斯特理工学院; 克莱姆森大学; 迈阿密大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SIPO通过对比式自我教师提供令牌级密集信用分配,统一强化学习与在线策略自蒸馏,在推理和代码生成基准上优于RLVR和OPSD基线。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)已成为提升大语言模型(LLMs)在各种任务上表现的标准范式,但其稀疏的结果奖励缺乏对中间步骤的令牌级信用分配。为解决这一问题,在线策略自蒸馏(OPSD)利用具有特权上下文的自我教师提供额外的密集学习信号。然而,由于自我教师往往过度自信,并对长推理轨迹施加过多惩罚,OPSD在实践中经常遇到困难。为缓解这一情况,我们提出自指导策略优化(SIPO),采用对比式自我教师提供密集信用分配。在每次迭代中,SIPO从当前策略中为每个提示采样多个轨迹,使用环境奖励对其进行评分,并通过将参考答案与组内错误配对,为每个轨迹构建两个教师上下文。然后,模型在这两种上下文中重新评估自身响应,利用两个教师对数概率之差作为令牌级反馈,使得两种上下文共享的偏差有望大幅抵消。由此产生的目标为每个轨迹提供令牌级优势:奖励仍设定每次更新的主要方向,而自我教师在令牌间重新分配信用。即使在所有轨迹均失败且组相对优势消失的组中,SIPO仍能提供学习信号。通过保留对任务奖励的直接优化,同时提供密集的令牌级反馈,该方法桥接了强化学习与在线策略自蒸馏。在多个推理和代码生成基准上的大量实验表明,SIPO在无需外部教师或额外生成的情况下,优于RLVR和OPSD基线。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.