RISE-RL:基于评分规则的开放式强化学习选择性探索
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究针对开放式任务中LLM的强化学习探索难题,提出RISE-RL方法,通过奖励过滤与策略支持塑造实现选择性内化,在多领域基准评测中显著优于现有方法。
中文摘要 AI 辅助
将大语言模型(LLM)用于开放式任务面临挑战,因为其响应需满足多维度标准,且不存在单一正确的生成轨迹。现有基于评分规则的强化学习(RL)方法将细粒度的标准级反馈压缩为标量奖励,在有限的在线探索下难以针对持续存在的能力差距。我们提出RISE-RL(Rubric-Informed Selective Exploration,基于评分规则的选择性探索),利用反复遗漏的评分规则标准引出仅通过无引导探索难以发现的特权轨迹。RISE-RL仅保留完整评分规则奖励超过自然rollout平均奖励的轨迹,随后在原始提示下重新评估这些轨迹,以强调自然策略仍弱支持的行为。所得引导信号通过单独的辅助目标优化,且在其额外益处消失后移除。针对4B和14B模型在写作、聊天、健康及科学领域的实验显示,RISE-RL在无引导评估下的所有评测基准中均取得最高平均得分。与标准Rubric-RL相比,其在4B规模下平均得分提升1.3分,在14B规模下提升3.3分,其中CreativeWriting-V3提升6.0分;同时还提升了创意写作的多样性,并在客观评分的医学和科学基准中取得增益。这些结果表明,通过奖励过滤和策略支持塑造进行选择性内化对开放式强化学习是有效的。
英文摘要
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration. We propose $\textbf{RISE-RL}$ (Rubric-Informed Selective Exploration), which uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. RISE-RL retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and then re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy. The resulting guidance signal is optimized through a separate auxiliary objective and removed once its additional benefit diminishes. Experiments with 4B and 14B models across writing, chat, health, and science show that RISE-RL achieves the highest mean score on every evaluated benchmark under guidance-free evaluation. Compared with standard Rubric-RL, it improves the average score by 1.3 points at the 4B scale and $\textbf{3.3 points at the 14B scale}$, including a $\textbf{6.0-point}$ gain on CreativeWriting-V3. It also improves creative-writing diversity and yields gains on objectively scored medical and scientific benchmarks. These results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.
发表机构
- Li Auto Inc.(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。