小型语言模型作为基于 rubric 的强化学习的评判者
Small Language Models as Judges for Rubric-Based Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究构建了两个基于 rubric 的评估数据集,对比三种提取标准级评判的方法,发现 Qwen3-1.7B 探针评判者表现最优,用作 GRPO 奖励模型时训练效率优于 8B 生成式评判者基线。
中文摘要 AI 辅助
基于 rubric 的强化学习将强化学习扩展到了具有精确答案或基于规则验证器之外的任务,它通过针对特定实例的标准对响应进行评分。然而,这使得奖励计算成本高昂:训练需要重复进行 rubric 评判,通常使用专有 API 或 7B 参数或更大的本地生成式 LLM 评判者。我们研究小型语言模型是否可以作为高效且可靠的基于 rubric 的评判者。为了使该问题可衡量,我们构建了 PointRubric 和 RaR-Science-Static 这两个基于逐点 rubric 的评估数据集,它们具有特定实例的标准和逐项满意度标签。我们比较了从小型模型中提取标准级评判的三种方法:生成式裁决、是/否对数概率边际和探针评判者。在两个数据集上,Qwen3-1.7B 探针评判者在这些方法中实现了最强的标准级一致性,优于生成式和对数概率评判者。将其用作 GRPO 奖励模型时,它在 RaR-Science rubric 分数上将策略从 0.232 训练到 0.643,而 8B 生成式评判者基线的分数为 0.594,同时基线需要多 10.7 倍的奖励评判时间。任务和领域迁移实验进一步表明,探针评判者在各种设置中保留了标准级奖励结构。
英文摘要
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific criteria. However, this makes reward computation expensive: training requires repeated rubric judging, often with proprietary APIs or local generative LLM judges with 7B parameters or more. We study whether smaller language models can serve as efficient and reliable rubric-based judges. To make this question measurable, we construct PointRubric and RaR-Science-Static, two pointwise rubric-based evaluation datasets with instance-specific criteria and itemwise satisfaction labels. We compare three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges. Across both datasets, the Qwen3-1.7B Probe judge achieves the strongest criterion-level agreement among these methods, outperforming Generative and Logprob judges. Used as a GRPO reward model, it trains a policy from 0.232 to 0.643 on RaR-Science rubric score, compared with 0.594 for an 8B Generative judge baseline, while the baseline requires 10.7$\times$ more reward-judge time. Task and domain transfer experiments further suggest that Probe judges preserve criterion-level reward structure across settings.
发表机构
- New York University(纽约大学)
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。