arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoRS:面向开放式强化学习的奖励系统在线策略自进化

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

Weiyuan Li, Aili Chen, Xintao Wang, Yikai Zhang, Qingqing Dong, Jinghan Xu, Hongru Hou, Wenxuan Zhao, Chengkun Lang, Jun Gao, Yuanli Guo, Hongcheng Guo, Yanghua Xiao, Deqing Yang

arXiv 2609.12459首次发表:更新:

发表机构

Fudan University; Shanghai Key Laboratory of Data Science; Nankai University; Hello Group(复旦大学; 上海市数据科学重点实验室; 南开大学; 挚文集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EvoRS提出一种自进化强化学习框架,通过智能体设计器从在线策略经验中动态更新奖励有向无环图,解决开放式任务中固定奖励系统因奖励黑客或区分度下降而失效的问题,在写作和角色扮演中显著提升质量。

AI 中文摘要

开放式强化学习通常依赖基于规则的奖励来处理没有直接可验证答案的任务。然而,策略与奖励系统形成了一个动态反馈回路:随着策略优化当前奖励,初始有效的奖励系统可能因奖励黑客攻击或响应区分度降低而变得不可靠。因此,奖励系统应在训练过程中进化而非保持固定。现有的动态规则方法调整评估标准,但奖励失败也可能源于评分机制或信号组合。我们提出EvoRS,一个自进化的强化学习框架,它从在线策略经验中进化奖励系统,将其表示为可执行的奖励有向无环图(Reward-DAG)。具体而言,一个智能体设计器从在线策略回放和奖励轨迹中更新该系统,以维持训练时的可靠性。在写作和角色扮演任务中,EvoRS在所有三种评判标准下均达到最佳质量,分别比策略高出2.107分和4.767分,同时减少了奖励黑客攻击和覆盖失败,并保持了奖励的信息量。消融实验证实,一个全面的固定奖励系统在开放式任务中无法保持可靠,必须在整个训练过程中进化。

英文摘要

Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal composition. We introduce EvoRS, a self-evolving RL framework that evolves the reward system from on-policy experience, representing it as an executable Reward-DAG. Specifically, an agentic designer updates this system from on-policy rollouts and reward traces to maintain train-time reliability. Across writing and roleplay, EvoRS achieves the best quality under all three judges, outperforming the policy by \(2.107\) and \(4.767\) points, respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations confirm that a comprehensive fixed reward system cannot remain reliable in open-ended tasks and must evolve throughout training.

Comments38 pages, 10 figures, 24 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑