发表机构
Amazon Web Services; University of Illinois Urbana-Champaign(亚马逊网络服务; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Cliff是一种利用现成LLM识别推理轨迹首次错误的奖励塑形策略,在12种场景下可提升推理性能,比策略蒸馏高15%、比标准GRPO高7%,为RLVR提供细粒度监督。
AI 中文摘要
可验证奖励强化学习(RLVR)已成为大语言模型(LLM)后训练的强大范式,但它依赖粗糙的结果奖励,导致对中间推理过程的指导有限。现有方法如过程奖励建模和策略蒸馏引入了额外约束,例如依赖专用奖励模型或假设教师与学生的推理模式相同。然而,我们观察到,一旦推理过程首次出错,评估后续推理提供的额外信息有限,因为它已基于无效前缀。因此,我们提出Cliff,一种奖励塑形策略,利用现成LLM作为教师识别每个 rollout(轨迹)中的首次错误,该轨迹自然分解为正确前缀和错误后缀两部分。Cliff将此信号转换为 token(令牌)级优势,为正确前缀分配正优势,为后续部分分配负反馈。在12种不同场景下的实验表明,Cliff始终提升推理性能,在教师能力一般的情况下,比策略蒸馏性能提升15%,比标准GRPO提升7%。此外,我们分析了Cliff中“真实值”的作用并研究其训练动态。这些结果表明,Cliff是一种简单、通用且有效的方法,可通过更丰富、细粒度的监督改进RLVR。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.