arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01548cs.LG

Range-GRPO:通过奖励区间间的成对关系进行策略优化

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM后训练中奖励不确定性未显式建模的问题,提出Range-GRPO半监督框架,通过成对比较保形校准的奖励区间进行策略优化,理论证明区间坍缩时退化为GRPO,实验显示在分布内外性能最优且资源消耗更少。

中文摘要 AI 辅助

随着大型语言模型(LLMs)的使用范围不断扩大,后训练对于使其适应下游任务变得越来越重要。然而,获取可靠的监督信号仍然代价高昂,尤其是在没有参考答案或可执行验证器的领域。LLM-as-a-Judge为未标注的响应提供了可扩展的伪奖励,但单一的分数并不能明确表示奖励的不确定性。这促使我们将伪奖励表示为经过保形校准的奖励区间。我们提出了Range-GRPO,一种半监督的后训练框架,它结合了有限的标注数据和未标注的提示。在群体相对策略优化(GRPO)中,学习信号依赖于每个滚动组内的相对奖励比较。所提出的目标函数成对比较奖励区间,而不是将其简化为点奖励,从而使区间不确定性能够影响这些信号的幅度和方向。我们的理论分析刻画了这一区别,并表明当所有奖励区间坍缩为点时,所提出的目标函数恢复为GRPO的优势函数。实验上,在评估的半监督方法中,Range-GRPO在分布内和分布外的平均性能上均达到最高,同时需要更少的训练资源。

英文摘要

As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr$.$GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.

发表机构

  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

↑