RL Forgets! Towards Continual Policy Optimization
强化学习会遗忘!迈向持续策略优化
机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ; Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education(教育部计算机网络和信息集成重点实验室(东南大学)) ; Zhongguancun Academy(中关村科学城公司) ; Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
专题命中 图文多模态 :multimodal(abstract)
AI总结 研究持续训练中强化学习易遗忘问题,引入MRCL基准测试,提出持续策略优化框架CPO,通过参数移动正则化减少遗忘,多模型规模实验验证了其有效性。