arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MInTRL:离策略干预可以提升在策略强化学习

MInTRL: Off-policy Intervention can boost On-policy RL

Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong

arXiv 2609.12419首次发表:更新:

发表机构

Amazon Web Services; Boston University(亚马逊云服务; 波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出最小干预强化学习(MInTRL),通过在在策略轨迹生成中引入稀疏局部干预,以序列级优势回归目标训练,在不牺牲可学习性下扩展探索,在数学和代码基准上优于在策略和离策略基线。

AI 中文摘要

具有可验证奖励的强化学习通常采用在策略(on-policy)方式进行,这使训练数据贴近当前策略,但将学习限制在策略自身能够发现的轨迹上。另一方面,诸如监督微调等离策略(off-policy)方法可以利用超出基础模型能力的外部知识,但可能遭受较大的分布偏移。因此,关键挑战在于在不牺牲可学习性的前提下扩大探索范围。在本工作中,我们引入了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),该方法通过在原本在策略的轨迹生成过程中进行稀疏的、局部的干预来扩展探索边界。在生成阶段,一个裁判干预策略(judge-intervention policy)定期审查当前策略的输出,用简短的修正替换错误的后续片段,并立即将控制权交还给策略。在训练阶段,MInTRL采用序列级优势回归目标(sequence-level advantage-regression objective),从而无需重要性采样(importance sampling)。我们表明,稀疏的局部干预可以显著提升覆盖范围,超越有限预算下的在策略采样,同时保持所得轨迹整体的在策略性质。在数学和代码基准测试中,MInTRL持续优于标准的在策略和离策略基线。消融实验表明,MInTRL在自我干预(self-intervention)以及不同裁判策略下均保持有效,而性能在中等干预强度时达到峰值,这凸显了最小干预的重要性。这些结果确立了最小干预作为增强在策略强化学习的有效范式。

英文摘要

Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's output, replaces erroneous suffixes with short corrections, and immediately returns control to the policy. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. We show that sparse, local interventions can substantially improve coverage beyond finite-budget on-policy sampling while preserving the overall on-policy nature of the resulting trajectories. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines. Ablations show that MInTRL remains effective with self-intervention and across different judge policies, while performance peaks at moderate intervention intensity, highlighting the importance of intervening minimally. These results establish minimal intervention as an effective paradigm for enhancing on-policy RL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑