arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23457cs.LGcs.AI

RLVR$^{2}$:基于可验证评分标准的排序强化学习

RLVR$^{2}$: Reinforcement Learning with Verifiable Rubric-based Ranking

Hao Li, Zhengkun Zhang, Gangqiang Hu, Zhen Zhang, Yude Gao, Dai Dai, Jing Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对多维评分标准下奖励聚合的不可靠性,提出RLVR$^2$,通过组内排序和潜在效用合并生成训练信号,在三个规模、16个基准上优于基线。

中文摘要 AI 辅助

基于可验证奖励的强化学习(RLVR)正从具有明确正确性信号的任务(如数学和代码)扩展到由多维评分标准所指定的多面质量要求。由于策略优化每次轨迹仅消耗一个标量,基于评分标准的流程必须将多个标准分数映射为标量奖励。这种聚合通常被视为分数缩放,但它隐含地决定了训练过程中质量维度之间的权衡方式。当前的主流做法是对每个标准进行归一化并取线性组合,这假设了各标准之间的基数分数差异具有可比性,并且一个标准上的收益可以补偿另一个标准上的失败;当标准在语义上异构时,这两个假设都不可靠。我们提出了基于可验证评分标准的排序强化学习(RLVR$^2$),这是一种用于基于评分标准的RLVR的可验证排序范式。对于每个标准,RLVR$^2$将评分标准分数转换为特定于标准的组内序数结果,从生成的比较矩阵中恢复潜在效用,并将这些效用合并为一个训练信号。通过仅保留组内排序并丢弃原始分数幅度,RLVR$^2$避免了校准异构评分标准尺度的问题。它还支持保持目标不变的属性调整:与观察到的排序相关但不是训练目标的辅助属性可以进入估计,而无需扩展评分标准或直接奖励它们。在三个模型规模和16个基准测试中,RLVR$^2$持续优于具有代表性的基于评分标准的基线,在大多数基准测试的每个规模上均取得了最佳整体性能。分析表明,它在保持质量目标的同时,控制了与推理效率和响应格式相关的系统性影响。

英文摘要

Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by multi-dimensional rubrics. Since policy optimization consumes one scalar per rollout, rubric-based pipelines must map multiple criterion scores into a scalar reward. This aggregation is often treated as score scaling, but it implicitly determines how quality dimensions trade off during training. The prevailing practice, normalizing each criterion and taking a linear combination, assumes that cardinal score differences are comparable across criteria and that gains on one criterion compensate for failures on another; both assumptions are unreliable when criteria are semantically heterogeneous. We propose Reinforcement Learning with Verifiable Rubric-based Ranking (RLVR$^2$), a verifiable ranking paradigm for rubric-based RLVR. For each criterion, RLVR$^2$ converts rubric scores into criterion-specific within-group ordinal outcomes, recovers a latent utility from the resulting comparison matrix, and merges these utilities into one training signal. By retaining only within-group ordering and discarding raw score magnitudes, RLVR$^2$ avoids calibrating heterogeneous rubric scales. It further supports objective-preserving attribute adjustment: auxiliary attributes that correlate with observed rankings but are not training objectives can enter the estimation without expanding the rubric or rewarding them directly. Across three model scales and 16 benchmarks, RLVR$^2$ consistently outperforms representative rubric-based baselines, achieving the best overall performance on most benchmarks at every scale. Analysis shows it controls systematic effects tied to reasoning efficiency and response formatting while preserving the quality objective.

发表机构

  • Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑