arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化学习会遗忘!迈向持续策略优化

RL Forgets! Towards Continual Policy Optimization

Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei

arXiv 2607.04364首次发表:更新:

发表机构

School of Computer Science and Engineering, Southeast University; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education; Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence(东南大学计算机科学与工程学院; 教育部计算机网络和信息集成重点实验室(东南大学); 中关村科学城公司; 中关村人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究持续训练中强化学习易遗忘问题,引入MRCL基准测试,提出持续策略优化框架CPO,通过参数移动正则化减少遗忘,多模型规模实验验证了其有效性。

AI 中文摘要

持续训练后适应不断发展的任务越来越重要,强化学习本应不易遗忘,但证据不足。为此引入MRCL基准测试,发现强化学习仍会严重遗忘。提出持续策略优化框架CPO,通过理论合理的参数移动正则化限制策略漂移,多实验表明其能减少遗忘并保留甚至提升预训练模型能力。

英文摘要

Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent studies report that reinforcement learning is less prone to forgetting than supervised fine-tuning, motivating the view that RL is inherently resistant to forgetting. However, this view remains insufficiently validated, as existing evidence is largely drawn from outdated or homogeneous benchmarks. We revisit this assumption by introducing MRCL, a Multimodal Reasoning Continual Learning benchmark built from recent and diverse multimodal reasoning tasks. Experiments on MRCL show that standard reinforcement learning still suffers from catastrophic forgetting during continual post-training. We trace the failure to an objective mismatch. The KL regularization used in common policy optimization methods is evaluated on current-task data, whereas forgetting is caused by behavioral drift on prior-task distributions. To address this problem, we propose Continual Policy Optimization (CPO), a replay-free method grounded in a prior-task behavioral KL objective. CPO derives a local Fisher surrogate from the historical KL objective and uses parameter movement as a gradient-free proxy for Fisher sensitivity, enabling sparse regularization with negligible additional overhead. Experiments on three model scales and comparisons with multiple RL baselines show that CPO consistently reduces forgetting while maintaining effective adaptation and preserving broader pretrained capabilities. The implementation code is available at https://github.com/MaolinLuo/CPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑