arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RankBuffer:面向开放式生成的基于排序的高效奖励

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Zixuan Yang, Yiqun Chen, Qi Liu, Wei Yang, Erhan Zhang, Liyi Chen, Qimeng Wang, Yan Gao, Jiaxin Mao

arXiv 2609.36652首次发表:更新:

发表机构

Renmin University of China; University of Southern California; Xiaohongshu Inc.(中国人民大学; 南加州大学; 小红书公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放式生成中逐点奖励难以校准的问题,提出RankBuffer方法,通过维护有序缓冲区并采用区间粗排与局部精排相结合的方式构建相对奖励,在四个基准上优于逐点基线且显著降低评判成本。

AI 中文摘要

开放式生成缺乏标准答案,这使得逐点奖励难以针对基于组的强化学习进行校准。直接对同一查询的生成结果进行排序提供了一种更合适的相对奖励信号,但现有的基于排序的奖励方法可能会产生可观的评判成本。我们引入了RankBuffer,它维护一个有序的、针对特定查询的先前已评判响应缓冲区,作为可复用的质量尺度。每个生成结果首先通过一次独立的粗略评判被插入到一个锚定区间内,之后只有被分配到同一区间的生成结果才进行局部精细排序。由此产生的完整顺序被转换为有界的排序奖励,同时边界扩展、局部细化和非活动锚点剪枝使缓冲区能够随着策略的演变而自适应。在四个开放式生成基准测试中,RankBuffer始终优于所有逐点基线方法。它还与最强的基于排序的奖励基线方法达到了几乎相当的性能,同时大幅降低了评判成本。消融实验证明了局部精细排序和锚点响应内容的重要性,而缓冲区分析表明,由生成结果衍生的锚点逐步扩展并细化了所覆盖的质量尺度。这些结果确立了响应复用作为构建高效相对奖励的有效方法。

英文摘要

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑