Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
利用群体回滚中的误差多样性进行强化学习
机构 * Peking University(北京大学) ; JD.COM(京东公司) ; Shanghai Innovation Institute(上海创新研究院)
AI总结 本文提出EDAS方法,通过利用群体回滚中的误差多样性来提升强化学习的效果,通过调整错误回滚的优势信号,鼓励模型保持多样化的推理路径,从而提高训练成功率。
Comments Code available at https://github.com/EDAS-jd/EDAS