发表机构
University of British Columbia; Microsoft(不列颠哥伦比亚大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对评分标准奖励聚合与评判成本问题,提出基于项目反应理论的RRT方法,用双参数模型和在线EM更新,在多个基准上超越GRPO并减少评判请求。
AI 中文摘要
许多语言任务没有可以自动检查的单一答案。评分标准为评判这些任务的回答提供了准则。对于强化学习,最终的评判结果必须合并为标量奖励。常见的方法是对满足的准则所分配的分数求和。这样,不同的评判模式可能获得相同的奖励,且固定的分数编码了每个准则应占多少权重,而非其评判结果对当前生成轨迹的区分强度。除了这个聚合问题外,评判完整的评分标准需要随着准则数量的增加而增加评判请求次数。为解决这些局限性,评分标准反应理论(RRT)在评分标准准则是共享目标的单调指标时,衡量质量并选择准则。RRT不是累加分配分数,而是使用双参数项目反应模型,将评判模式视为关于特定于评分标准的标量质量的证据。在该模型下,其似然分数最大化质量的局部信噪比。其反应参数网络(RPN)读取提示和准则文本以预测准则难度和区分度。随着训练期间策略分布的变化,RRT使用在线期望最大化从当前生成轨迹的评判结果更新RPN。以Qwen3.5-4B作为策略,RRT在Medical、Science、Rubrics as Rewards Science和RubricBench上的宏准则分数比组相对策略优化(GRPO)高出1.7分。在Medical和Science中的困难和非常困难准则上,RRT比GRPO高出2.8至5.6分。在准则预算减半的情况下,使用冻结的RPN进行自适应Fisher选择,四个数据集的宏准则分数保持在GRPO完整评判的0.1分以内。这些结果表明,RRT可以减少评判请求次数,同时保持与GRPO相当的竞争力。
英文摘要
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.