持续推理环境:诊断与利用持续RLVR中的共享推理
Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR
- State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室(BIGAI))
- Beijing Institute of Technology(北京理工大学)
- Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
- State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室、北京大学智能科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出持续推理环境,针对持续RLVR中顺序训练性能低于MTRL的问题,提出CPR方法利用共享推理,使模型平均达到MTRL级性能。
中文摘要 AI 辅助
带可验证奖励的强化学习(RLVR)通常会在多个任务上对推理模型进行后训练,而随着新任务加入重新运行多任务RLVR(MTRL)会导致能力扩展成本高昂。因此我们研究持续RLVR,即每个任务到达时更新现有模型,核心问题是如此更新的模型能否达到联合训练模型的性能。为回答该问题,我们提出持续推理环境(Continual Reasoning Gym),一种将文本与视觉推理任务组织为五个任务序列的持续RLVR环境。在该设置下,我们得出两个关键观察:顺序RLVR存在适度遗忘,但最终性能仍低于MTRL。为解释后者,我们分解最终性能,发现遗忘仅占性能差距的一部分;为解释前者,我们识别出共享推理:可迁移的推理结构使得对一个任务的训练平均上能支持其他任务。因此我们提出持续提示重放(CPR),该方法通过重放先前任务的提示并利用当前策略重新生成其响应,来利用共享推理提升当前及未来任务的学习。平均而言,仅CPR能达到MTRL级别的性能。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.