发表机构
ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于评分标准的强化学习中响应内信用分配问题,提出CoRT方法,利用反事实重放对响应重新评分,将对比映射为权重来重新分配优势,实验表明该方法能有效提升训练效果,兼具简单性与稳定性。
AI 中文摘要
基于评分标准的强化学习通过根据明确标准评估模型输出丰富语言模型训练。在GRPO风格管道中,结构化判断被简化为标量响应级奖励并转换为响应级优势,均匀广播到所有生成令牌。这使得响应内没有明确的信用分配机制。我们提出CoRT,一种用于评分标准条件下GRPO的令牌级信用加权方法。CoRT使用反事实重放对原始评分标准条件提示和匹配的无标准提示下的相同采样响应重新评分。实验表明CoRT在绝大多数比较中优于匹配的响应级GRPO,平均增益4.4个百分点,该方法与基于学习的令牌级信用基线竞争,避免单独的相关性学习阶段。
英文摘要
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.