arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPD Before RL:用在线策略蒸馏为基于量规的强化学习进行热启动

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak, Yun He, Richard Yuanzhe Pang

arXiv 2610.02781首次发表:更新:

发表机构

New York University; Meta(纽约大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于量规的强化学习无法定位具体决策贡献的问题,提出先用量规特权在线策略蒸馏提供密集令牌监督,再优化量规奖励的两阶段框架,在健康与科学任务上取得最优性能且奖励黑客迹象有限。

AI 中文摘要

许多有用的语言模型任务无法通过精确的结果验证来评估。基于量规的强化学习(RL)通过根据明确标准对开放式回答进行评分来解决此问题。然而,由于奖励是在完整回答之后分配的,训练信号无法直接指出哪些个别决策对最终分数有贡献。我们提出一个两阶段训练框架,首先将量规用作特权教师上下文以提供密集的令牌级监督,然后将其用作进一步强化学习的奖励。在第一阶段,量规特权在线策略蒸馏(RP-OPD)中,无法访问量规的学生模型在学生生成的序列前缀处匹配知晓量规的教师模型的下一令牌分布。在第二阶段,强化学习直接优化量规奖励,并超越观察到的蒸馏平台期。我们在健康与科学任务上使用开放权重模型评估该框架。在HealthBench、ResearchQA和RubricHub Science上,我们比较了后训练方法,并改变强化学习前的SFT或RP-OPD训练量,发现我们的两阶段框架在评估的方法中取得了最高分数。RP-OPD + RL在RubricHub Science上显示出有限的奖励黑客迹象,而SFT + RL基线则越来越多地因声称符合量规但未提供所需内容而获得高奖励。这些发现支持在应用基于量规的强化学习之前,使用量规指导在线策略蒸馏。

英文摘要

Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑